What a DOI actually is

A URL is a location. A DOI is an identity. That distinction sounds small until a journal moves to a new publisher, a university redesigns its website, and suddenly every link in every reference list for the past decade returns a 404. The Digital Object Identifier was built to prevent exactly that.

A DOI is a persistent identifier: a short, standardised string — something like 10.1038/nature12373 — that names a piece of content rather than describing where it currently lives. The part before the slash is a prefix identifying the registrant (a publisher, a repository, a data archive). The part after is a suffix they assign to the specific item. Together they are unique, permanent and resolvable: paste a DOI into a browser prefixed with https://doi.org/ and a lookup service called the DOI resolver forwards you to wherever the content currently lives. When the content moves, the publisher updates that forwarding record. The DOI itself never changes.

The infrastructure behind this is maintained by the International DOI Foundation, and the most widely used registration agency for scholarly content is Crossref, which handles the majority of journal articles. Repositories and data archives often work through DataCite, which specialises in research datasets, software and other non-article outputs. Both organisations issue prefixes, maintain the resolver and set standards — but neither stores the content itself.

Why it matters for researchers and readers

If you have ever clicked a reference in a PDF only to land on a publisher's homepage — or nowhere at all — you have met link rot. Studies examining reference lists over time have consistently found that a substantial proportion of cited URLs eventually break. DOIs are the discipline's answer: a citable, machine-readable handle that outlasts the web address it points to.

For anyone looking for open access copies of research, the DOI is also the surest starting point. Tools like Unpaywall use the DOI to check whether a legal free version of a paper exists — a preprint on arXiv, a manuscript deposited in PubMed Central, a version uploaded to Zenodo. Without the DOI as a stable identifier, that matching would be far less reliable.

For researchers depositing their own work, repositories generate DOIs automatically. Zenodo, for instance, mints a fresh DOI for every upload, meaning a dataset or a conference poster becomes citable the moment it is deposited — even if it never appears in a journal. Creative Commons licences travel alongside the metadata, so the terms of reuse are attached to the same persistent record. DOAJ uses DOIs as part of the quality criteria it checks when evaluating whether a journal meets its standards for genuine open access.

A few things a DOI cannot do

Persistence depends on institutions keeping their promises. If a publisher ceases operating and no one maintains the resolver record, the DOI still exists but resolves to nothing. Preservation archives like Portico and CLOCKSS exist partly to guard against this — but they are separate systems, and a DOI alone is not a guarantee of perpetual access.

DOIs also do not validate quality. A paper in a predatory journal carries a DOI just as readily as one in a rigorously peer-reviewed venue. The identifier is a locator and a citation anchor, not a mark of credibility.