The archive you've probably already used without knowing it

Somewhere between the paper you found for free last Tuesday and the dataset cited in a Nature Methods article is a piece of infrastructure most researchers never think about — and most students have never heard named. It's called a digital repository, and it is, in the plainest sense, an organised online store for scholarly content: papers, datasets, theses, images, code, reports, and more. Not a journal. Not a search engine. Not a publisher. A repository holds things and makes them findable, often permanently, often for free.

The word "repository" sounds institutional and faintly dusty, which is ironic given how alive these systems are. Every day, researchers deposit new preprints to arXiv, upload datasets to Zenodo, and submit theses to institutional archives that feed into national and international discovery networks. If you have ever clicked a link to a PDF that landed somewhere other than a journal website — a university domain, a page ending in .org, a portal with an unfamiliar name — there is a reasonable chance you landed in a repository.

What they hold, and what they do with it

The content of repositories ranges widely. Subject repositories like arXiv (physics, mathematics, computer science, quantitative biology and related fields) and PubMed Central (biomedical and life sciences) hold preprints and peer-reviewed manuscripts — the intellectual output of research communities organised by discipline. Institutional repositories, run by universities and research organisations, tend to hold everything produced under one roof: journal articles deposited by faculty, doctoral theses, conference papers, technical reports, and increasingly research data. General-purpose repositories like Zenodo, operated by CERN and supported by the European Commission, will accept almost anything — a dataset, a piece of software, a poster, a short video — and assign it a persistent identifier so it can be cited.

That persistent identifier matters more than it might seem. Most repositories assign a DOI to every item deposited, which means that even if the hosting server changes or the URL is restructured, the record stays findable. A repository without persistent identifiers is a filing cabinet. One with them is a living part of the scholarly record.

Beyond storage, repositories do several things at once. They make metadata — author names, abstracts, keywords, dates, subject classifications — readable by search engines and harvesting services, so that a deposit in a university's own archive can surface in Google Scholar, BASE (Bielefeld Academic Search Engine), or OpenAIRE. They apply licence information, so a reader knows immediately whether they can reuse, translate, or adapt a work. Many use Creative Commons licences for this purpose, from the permissive CC BY (attribution only) to the more restrictive CC BY-NC-ND (no commercial use, no derivatives). And they preserve: a good repository commits to keeping its content available over time, which distinguishes it from posting a PDF to a personal webpage and hoping no one moves the server.

Why they exist, and who runs them

Repositories emerged from a simple frustration: scholarly knowledge was produced with public and institutional money but locked behind journal paywalls that most of the world — and many researchers — could not afford. The early internet offered an obvious workaround. Physicists began sharing preprints informally in the early 1990s; arXiv was founded in 1991 to formalise this, giving the community a stable home for manuscripts before and after peer review. The model spread. Librarians at universities built institutional repositories in the 2000s as part of a broader open-access movement, creating local infrastructure that could capture the research output of a campus. Funders and governments began mandating deposit — requiring that research they paid for be made publicly available — which accelerated adoption significantly.

Today, repositories are run by a range of organisations. Subject repositories are often run by learned societies, libraries, or research institutions with domain expertise. Zenodo is run by CERN. PubMed Central is operated by the US National Library of Medicine, part of the National Institutes of Health. Institutional repositories are typically managed by university libraries. The Directory of Open Access Repositories (OpenDOAR) tracks thousands of repositories worldwide, and the Directory of Open Access Journals (DOAJ) performs a related function for open-access journals — vetting them for quality and listing those that meet its standards.

Funding models vary. Most repositories are not-for-profit, sustained by institutional budgets, government funding, grants, or community membership. Zenodo, for instance, is free to use and does not charge for deposits or access. This is one of the things that distinguishes repositories from journals: they are not in the business of selling content; they are in the business of preserving and sharing it.

The relationship with open access

Repositories are one of the two main routes to open access. The other is publishing directly in an open-access journal. When a researcher deposits a version of their article — a preprint before peer review, or an accepted manuscript after — in a repository and makes it publicly available, that is called green open access. It costs nothing and works even when the final published version sits behind a paywall, because the version in the repository is free to read. Many funders and institutions now require this as a condition of grant funding, and most major publishers have policies — however complicated — that permit some version of an article to be self-archived.

This is worth pausing on. A repository deposit does not replace journal publication; in most disciplines, peer review and formal publication still matter enormously for career progression and academic credibility. What it does is separate access from publication: the journal can still exist, peer review can still happen, and the record of scholarship can still reach anyone with an internet connection. For researchers in low-income countries, for journalists, for clinicians in under-resourced settings, for students at institutions without comprehensive journal subscriptions, the repository version is often the only version available.

Versions, and why they differ

One genuine complexity worth understanding is the version question. A paper typically exists in several forms before and after publication: the original submitted manuscript, the version revised after peer review (the accepted manuscript), and the final typeset version of record produced by the publisher. Repositories usually hold one of the first two — often the accepted manuscript — because publishers generally retain copyright over the typeset version and restrict redistribution. This means the paper you find in a repository may look different from the journal version, and in rare cases may contain minor differences in wording or figures. It is the same intellectual content, produced by the same authors, peer-reviewed in the same process — just dressed differently.

Understanding this helps when you cite: where possible, cite the version of record using the journal DOI, and note the repository version if that is what you accessed. Many repositories now display both, making this easier than it used to be.

What this means for you

If you are a student, a researcher, or simply curious: repositories are not a workaround or a grey area. They are a designed, deliberate part of how scholarship moves through the world. Using them is legal, free, and often the fastest route to the full text of a paper. Finding free, legal copies of papers you need is a practical skill, and repositories — arXiv, PubMed Central, Zenodo, your own university's institutional repository, or a subject-specific one in your field — are where to look first.

If you are a researcher: depositing your work is one of the most effective things you can do for its reach. A paper accessible to anyone is a paper that can be cited, used, taught with, and built on. Repository deposit takes minutes; its effects are measured in years of access.