Guides

How to verify the citations in an AI-generated report

To verify a citation you have to answer four separate questions, and only one of them is about whether the reference is real. Does the cited work exist. Does its text support the sentence that cites it. Is every quoted span intact. Has the work been retracted or superseded. The first and the fourth are mechanical and cheap. The third is string matching against the fetched source body. The second needs somebody or something to read the source, and it is the one that fails most often on documents that look fine. This page gives a manual procedure for each, the order to automate them in, and an honest account of what automation still gets wrong.

Four failure classes, each with an invented example

Every example below is fabricated to show the shape of the failure. No real paper, author, journal, DOI or figure appears in any of them.

Class 1, the fabricated reference. The document cites “Ostrowski, K. and Vel, M. (2021). Thermal retrofit uptake in mid-rise housing stock. Journal of Building Performance 14(3), 221 to 238.” The journal is plausible, the volume and page range are plausible, the authors are plausible, and nothing about the reference exists. The tell is mechanical: the DOI does not resolve, or it resolves to an unrelated paper, and the title in quotation marks returns nothing. This class is the easiest to catch and the most talked about, which is why it is not where your attention should mostly go.

Class 2, the real reference that does not support the sentence. The document says “Retrofit grants cut household energy use by a fifth within two years [4].” Source 4 is a real, well-cited paper, correctly attributed, and what it reports is a modelled savings potential of about a fifth over a ten-year horizon in one city. The number matches. The unit matches. The modality does not: modelled potential is not measured outcome, and ten years is not two. A metadata check passes this citation with full marks. This is the most common failure on a competent document and the reason the other three classes are not enough.

Class 3, the altered or invented quote. The document writes that the authors concluded “uptake collapses without upfront subsidy.” The source says “uptake slows markedly in the absence of upfront subsidy in the two cohorts we observed.” The paper is real, the citation is correct, the sentiment is roughly preserved, and the quotation is not a quotation. A hedge was removed, the scope condition was dropped, and a verb was strengthened. The variant that is easier to catch is the quote that appears nowhere in the source at all.

Class 4, the retracted or superseded source. The document cites a 2019 panel study, invented for this example, that was retracted in 2024 after an error in its model specification came to light. The DOI resolves, the metadata is perfect, the paper is still downloadable, and it is still cited thousands of times. Superseding is the softer version of the same problem: the document cites a preprint whose published version revised the headline figure, and the citation points at the number that no longer stands.

The manual procedure

This is what a careful editor does, and it is worth doing by hand on ten citations before you automate anything, because it tells you which class your documents actually suffer from.

  1. Resolve the reference. Try the DOI at https://doi.org/<doi>. If there is no DOI, query Crossref or OpenAlex by title, and search the exact title in quotation marks. A reference that will not resolve three ways is class 1.
  2. Open the source and get the real text. The abstract is not the source. A claim about a subgroup, a robustness check or a limitation lives in the body, and comparing a sentence against an abstract produces confident nonsense in both directions.
  3. Locate the passage. Search the source for the sentence’s numbers first, then for its two most distinctive nouns. If the number is not in the source at all, stop and record it; you are done with this citation.
  4. Compare the claim, field by field. Number, unit, time period, population, direction, and modality. Modality is the field people skip: measured, modelled, projected, self-reported and simulated are five different claims and they get written as one.
  5. Check the quotation character by character. Any span in quotation marks has to appear in the source, allowing for whitespace, hyphenation and ellipses and nothing else. A paraphrase inside quotation marks is class 3 even when it is faithful.
  6. Check retraction and update status. Look up the DOI’s update records. Crossref publishes the Retraction Watch database it acquired from the Center for Scientific Integrity, and it carries retractions, corrections and expressions of concern. Retraction-aware research has the exact request shape and the reasons one signal is not enough.
  7. Write down the verdict and the span you used. A verdict with no quoted evidence behind it is an opinion, and in a month you will not remember which passage you read.

Budget five to fifteen minutes per citation when the full text opens cleanly, and considerably more when it does not.

What to automate, and in what order

Automate in ascending order of cost, because the cheap checks remove work from the expensive one.

  1. Resolve every reference. One API call each, no model involved. Catches class 1 completely, and catches nothing else.
  2. Sweep every DOI for retractions and updates. One lookup each, no model involved. Catches class 4. Do this before you spend anything on reading, because a retracted source’s support is moot.
  3. Match every quoted span against the fetched source body. String matching with normalisation for whitespace, hyphenation and ellipses. Catches class 3 without a model call. Report near misses with the closest span and an edit distance, because a near miss is usually a real alteration and a total miss is usually a fabrication.
  4. Bind each citation marker to the sentence carrying it, then triage mechanically. A pair whose numbers and a long word-overlap window already appear in the cited source’s extracted claims can pass without a model. Unresolvable markers come out of this step as findings.
  5. Send the remainder to a model, prioritised. Check every sentence carrying a strong claim first, such as a percentage, a currency amount, a magnitude word or a change verb. Then sample the rest. Quantitative sentences must never be waved through on word overlap, because that is exactly where class 2 lives.
  6. Cluster the sources for independence, last. Five outlets carrying one press release are one voice, and this changes how you read a well-cited document without changing any single verdict.

Steps 1 to 4 are cheap enough to run on everything. Step 5 is where the money goes, which is why any honest system reports how many pairs it actually read. What a verification receipt contains documents those counters field by field.

Where reference checkers stop

Most citation tools answer question one. They take a reference list, resolve the identifier, compare the metadata against Crossref, PubMed, arXiv or OpenAlex, and tell you whether the work is real and whether author, year and title agree. That is a genuinely useful check, it is fast, it is cheap, and it is the right tool if your documents fail at class 1.

It cannot see class 2, because class 2 is invisible in metadata. It cannot see class 3, because quotations are not in the metadata record. It sees class 4 only if it queries update records, which many do not. Support checking begins where the metadata record ends: at the source’s body text, bound to one sentence of your document at a time.

Reading a verdict table

A per-sentence checker gives you one row per citation-sentence pair. The vocabulary Hyperresearch, the AI deep research API, uses is supported, partially-supported, unsupported and wrong-source, and the useful part of each row is the evidence span, the text from the source the checker decided on, because that is what lets you overrule it.

Work the table in this order.

wrong-source first. The claim may well be true and the pointer is wrong, which usually means a renumbering error, a shifted reference list, or a model attributing a figure to the nearest citation. One of these often implies several, so check the neighbours of every hit.

unsupported next. Read the evidence span before you touch the sentence. Roughly speaking, either the source does not contain the claim, in which case the sentence is wrong, or the checker never saw the right part of the source, in which case the pipeline is wrong. These need different fixes and look identical in the table.

partially-supported last, and expect most of your rows here on a document that was written carefully. It usually means one sentence is making two claims and the citation covers one of them. Splitting the sentence, or moving the citation to the clause it actually supports, resolves most of these without changing a single fact.

Then read the counters, not just the rows. How many pairs did the document contain, how many passed mechanically, how many reached a model. A table of forty clean verdicts on a document with three hundred citations is not a clean document.

False positives, and what a checker cannot know

Automated support checking is wrong in both directions, and a tool that does not tell you how is asking for more trust than it has earned.

Five things produce a false unsupported:

  • The full text was never retrieved. If the checker compared your sentence against an abstract or a landing page, it will report absent support for support that is present on page 11. This is the single largest cause, and it is a fetching failure wearing a verification failure’s clothes.
  • The support is in a figure or a table published as an image. Text-only verification cannot read a chart. A claim resting on one is out of scope, not unsupported.
  • The claim is a correct paraphrase with no shared vocabulary. A sentence can restate a source accurately in words the source never uses.
  • The claim aggregates two sources and cites one. This is a citation-placement problem the author should fix, but the sentence is not false and the flag reads as though it were.
  • The source changed. A page that is dynamic, paginated, or updated between the run and the check is a different source from the one the writer read.

And three things produce a false supported:

  • The source contains the number in a different context, and the numbers matching was enough.
  • The source is quoting somebody else, including quoting the very claim it goes on to reject.
  • The source says it, and the source is wrong. A supported verdict is a claim about the source, never a claim about the world. No citation checker is a fact checker, and one that implies otherwise is misleading you.

Two limits are structural rather than statistical. Citation extraction covers numbered [N] and wikilink [[note-id]] markers; author-year and footnote styles need a resolver and are not supported today. And sampling means what it says: pairs outside the sample were not read, and the only defence is a system that tells you the count.

Finally, the number you should ask every vendor for, including this one. Hyperresearch has not yet measured its cite-checker’s false-positive rate on a defined sample, so this page does not quote one. Until it does, treat a verdict table as a triage queue for a human reader rather than a pass mark, and treat any vendor quoting a precision figure without publishing the sample and the method the same way.

You can run the four checks on your own text without an account at /verify, or over HTTP through the citation verification API.

By Jordan Gibbs · Updated 2026-09-21