Guides

How to check whether a paper is retracted, programmatically

To check whether a paper is retracted, ask Crossref for its DOI and read the updated-by array. Or download the Retraction Watch CSV and join on OriginalPaperDOI. Both are free and need no key. The trap is the field name. Crossref’s own documentation points at update-to, and on a retracted paper that field often holds no retraction at all.

Where the data comes from

Retraction Watch, run by the Center for Scientific Integrity, built the largest curated database of retraction notices. On 12 September 2023 Crossref announced it had acquired the database and would make it public. Crossref paid a USD $175,000 fee and pays “USD $120,000 each year, increasing by 5% each year”. The Retraction Watch blog stays separate from Crossref.

Crossref’s documentation describes the data this way:

In September 2023, we acquired the Retraction Watch database from the Center of Scientific Integrity and have made it publicly available. The database contains retractions gathered from publisher websites and is updated every working day by Retraction Watch. Some other update types, such as expressions of concern and corrections, are also included in the data, but these are not as comprehensive as retractions.

Take the last sentence at its word. The retraction data is thorough. The expression-of-concern and correction data is incomplete.

Crossref states no licence for the dataset, and GitLab reports none for the repository. Query it freely. Don’t assume you may redistribute it.

Read updated-by, not update-to

Crossref’s documentation says “Retractions are included in the update-to field of json files in the REST API”. That is true of the retraction notice. The notice’s update-to points at the paper it retracts. The retracted paper itself carries updated-by, which points at the notices about it.

A real record shows the difference. https://api.crossref.org/works/10.1038/s41586-020-2801-z is a retracted Nature paper on room-temperature superconductivity. Read on 2026-09-27, its update-to array holds two entries, and neither is a retraction (fields trimmed):

[
{ "DOI": "10.1038/s41586-020-2801-z", "type": "expression_of_concern",
"source": "retraction-watch", "updated": { "date-time": "2022-02-15T00:00:00Z" } },
{ "DOI": "10.1038/s41586-020-2801-z", "type": "expression_of_concern",
"source": "retraction-watch", "updated": { "date-time": "2021-08-30T00:00:00Z" } }
]

Its updated-by array has the retraction twice, once from Retraction Watch and once from the publisher (retraction entries only):

[
{ "DOI": "10.1038/s41586-022-05294-9", "type": "retraction", "label": "Retraction",
"source": "retraction-watch", "updated": { "date-time": "2022-09-26T00:00:00Z" },
"record-id": "39363" },
{ "DOI": "10.1038/s41586-022-05294-9", "type": "retraction", "label": "Retraction",
"source": "publisher", "updated": { "date-time": "2022-09-26T00:00:00Z" } }
]

Code that reads update-to concludes this paper is fine. Code that reads updated-by gets the retraction, the notice DOI, the date, and a record-id that joins to the CSV. The 1998 Lancet paper by Wakefield and colleagues, retracted in 2010, behaves the same way: update-to is null and the retraction sits in updated-by.

Two more details. source is either publisher or retraction-watch, and only Retraction Watch entries carry a record-id. And the full updated-by array on the Nature paper is a history: a correction on 2020-11-20, expressions of concern on 2021-08-30 and 2022-02-15, then the retraction on 2022-09-26. An answer cached in 2021 would have been right then and wrong now.

Free ways to get the answer

One DOI, live. Identify yourself with mailto, then read updated-by:

Terminal window
curl -sS \
-H 'User-Agent: YourTool/1.0 (https://example.org; mailto:you@example.org)' \
'https://api.crossref.org/works/10.1038/s41586-020-2801-z' \
| jq '.message["updated-by"]'

The literal mailto: token matters. With it, or with a mailto= query parameter, the response header x-api-pool reads polite-single and the rate limit is 10 requests a second. A bare email address in the User-Agent gets public-single and 5 a second.

Every notice that points at a DOI. filter=updates:<DOI> returns the notice records themselves:

Terminal window
curl -sS 'https://api.crossref.org/works?filter=updates:10.1038/s41586-020-2801-z&rows=5&mailto=you@example.org'

For the Nature paper this returns three records: the retraction note, a publisher correction, and the paper itself. The paper appears because Retraction Watch records its expressions of concern as updates pointing at the paper’s own DOI.

The whole database as a CSV. Crossref publishes it as a git repository:

Terminal window
git clone https://gitlab.com/crossref/retraction-watch-data

The repository holds two files, README.md and retraction_watch.csv, and gets a new commit most working days. The columns begin Record ID,Title,Subject,Institution,Journal,Publisher,Country,Author,URLS,ArticleType,RetractionDate,RetractionDOI,RetractionPubMedID,OriginalPaperDate,OriginalPaperDOI,OriginalPaperPubMedID,RetractionNature,Reason,Paywalled,Notes. On 2026-09-27 the file held 72,684 records in about 67 MB. Join on OriginalPaperDOI, and lowercase both sides first: the CSV keeps the publisher’s DOI case and the API lowercases.

Crossref Labs serves the same file at https://api.labs.crossref.org/data/retractionwatch, with Windows line endings. It has no per-DOI lookup. Adding ?doi= returns the whole file.

The update-type filter finds notices, not retracted papers

filter=update-type:retraction returned 75,785 records on 2026-09-27. Those records are the notices. Filter the 1998 Lancet paper’s DOI the same way and you get zero results, though the paper is retracted.

The filter also accepts anything. filter=update-type:zzzznotreal returns "status": "ok" and zero results, so a typo looks like a clean record.

The values themselves are messy. Crossref’s Crossmark documentation defines twelve update types: addendum, clarification, correction, corrigendum, erratum, expression_of_concern, new_edition, new_version, partial_retraction, removal, retraction and withdrawal. The live data doesn’t stick to them. A facet query on 2026-09-27 returned 34 distinct values, from correction at 214,535 down to one-offs such as err, Retraction with a capital R, expression-of-concern with hyphens, retration and 68818. Lowercase the value, and treat anything outside the twelve as unknown rather than as a clean record.

Retraction, expression of concern and correction mean different things

Crossref doesn’t define these terms. Its guidance on corrections and retractions lists COPE’s retraction guidelines and NISO’s recommendations as reading. COPE’s Retraction guidelines, Version 3, August 2025, give the purpose:

Retraction is a mechanism for correcting the literature and alerting readers to articles that have such seriously flawed or erroneous content or data that their findings and conclusions cannot be relied upon.

And the test: “Editors can decide to retract a publication if they no longer have confidence in the results and conclusions reported in the paper.”

COPE’s separate guideline on expressions of concern says one “might be warranted if there are major and credible concerns about the reliability of a publication which do not meet the criteria for a retraction, or conclusive evidence cannot be obtained for some time.”

A correction is the other outcome. COPE says to correct rather than retract when “correction would sufficiently deal with the errors or concerns raised, provided that the main results and conclusions are not unduly affected by the correction”.

So a correction says the conclusions stand. An expression of concern says nobody knows yet. A retraction says they don’t stand. A pipeline should treat each one differently, and a single is_retracted boolean loses that. COPE’s guidelines are licensed CC BY-NC-ND, so quote them rather than paraphrase.

AI research tools miss retractions

A study measured this. Labenbacher and colleagues, “Performance of AI Tools in Citing Retracted Literature : Content Analysis”, Journal of Medical Internet Research volume 28, 1 May 2026, DOI 10.2196/88766, tested nine free AI tools. They picked 15 retracted articles from the Retraction Watch database, asked each tool five standard questions about each article, asked every question twice, and had two researchers rate the answers. Their conclusion:

Freely available GenAI tools are currently not able to detect, exclude, or appropriately flag retracted scientific literature. The widespread and confident reproduction of retracted studies represents a substantial threat to research integrity, particularly in medical and evidence-based fields. Until retraction-aware verification mechanisms are systematically integrated, independent source checking remains essential when using AI-assisted literature tools.

No tool answered every question correctly. The three research-focused tools “failed to produce a single fully correct response set”. Retracted articles “were frequently included in topic overviews without warning, with error rates exceeding 40% in several tools”.

The cause is simple. A retracted paper still resolves, keeps its metadata, stays downloadable, and keeps collecting citations. Fetching it looks like fetching any other paper. A tool only catches a retraction if it looks the DOI up in a register, as a separate step.

How Hyperresearch checks retractions

Hyperresearch, the AI deep research API, checks retractions in two places.

While it reads sources. The model that summarises each fetched page can mark it retracted, for example when the page carries a retraction banner. A marked note gets a retracted tag and a much lower quality score, and the run’s coverage notes list it as retracted. That is a reading of the page, not a registry lookup.

Before a report ships. The sweep collects every DOI in the report text and the stored DOI of every source the report cites, including sources the report cites by link without printing a DOI. On the verification API it also takes the DOIs in the dois array and on any supplied source. DOIs with brackets, such as the Lancet paper’s 10.1016/S0140-6736(97)11096-0, are kept whole.

Each DOI is looked up live. The sweep asks OpenAlex for its is_retracted flag, then reads the paper’s own Crossref record, taking notices from updated-by as described above. Crossref calls carry a mailto contact, so they use the polite pool. Each DOI gets one status:

  • retracted, from a retraction, withdrawal or removal notice, or from OpenAlex’s flag.
  • expression-of-concern, from an expression of concern.
  • corrected, from a correction, erratum or corrigendum.
  • ok, when neither register holds any of those notices.
  • unavailable, when a register errored or rate-limited the request. That means the answer is unknown, not clean.

When a paper has several notices, the most serious one wins. The row carries that notice’s date and a link to it when Crossref has the notice.

A retracted DOI fails the check unless the report itself says the paper was retracted. The standalone verification API looks for the word “retract” in the sentence around the DOI. Inside a run, the check looks within 200 characters of the citation. The exception exists because a report about research integrity has to cite retracted papers. When a run fails this check it doesn’t ship. It stops in a blocked state until a person looks at it. On the verification API, the report comes back failed.

Expressions of concern, corrections and unavailable lookups don’t stop a run. They are counted and listed on the receipt, so you can see which DOIs were not confirmed. On the verification API, an unavailable lookup makes the result partial rather than passed.

Two limits apply. Only DOIs are checked, so a working paper, a report or a web page without one can’t be checked against any register. And a DOI that neither OpenAlex nor Crossref has a record of comes back ok, because there is no notice to find. Expressions of concern and corrections are also only as complete as the registers, which Crossref says are less thorough for those than for retractions.

Every DOI’s result lands in the run’s receipt, so you can audit the sweep yourself. What a verification receipt contains documents the retraction_sweep table field by field. Verifying citations in AI-generated text covers the other checks. To sweep a bibliography you already have, send up to 100 DOIs in the dois array of the citation verification API. It needs no document.

By Jordan Gibbs · Updated 2026-09-27