Product

A verification receipt is the record of what was checked

A verification receipt is the machine-readable record of every check run against a report’s citations before that report shipped. It is one JSON object with a pass or fail flag, a ship-gate checklist, and four tables: a cite-check verdict for each citation and the sentence it supports, a quote-integrity row for each quoted span, a retraction row for each DOI, and an independence audit that groups sources that are really one voice. Every Hyperresearch run produces one. The standalone verification endpoints produce the same object for a document Hyperresearch did not write.

The receipt is one object with six parts

The receipt schema is fixed and versioned. At the top level it carries id, passed, created_at, the engine version under engine.hyperresearch, a billing block with llm_pairs and price_cents, and then the checklist and the four tables.

passed is a boolean and it is the ship gate. A report whose receipt says passed: false did not clear the gate.

checks[] is the ship-gate checklist. Each row has a name, an ok boolean, a detail string that says in words what the check found, and an outcome of passed, warning, failed or skipped. The names are the real internal gate names, not marketing labels: quote-integrity, retracted-citations, cite-check-resolved, code-attributed, no-scaffold-leak, and the artifact-presence checks for each step that was supposed to run. Some rows are informational by design and carry a warning outcome without setting ok: false, so a stylistic miss cannot flip a run’s passed flag.

Then the four tables, each with its own counts and its own rows[].

Cite-check binds every citation to the sentence that used it

cite_check reports these counters:

  • pairs: how many citation-sentence pairs the document contained.
  • auto_passed: how many cleared mechanically without a model call.
  • dangling: citation markers with no resolvable source.
  • sampled_for_llm: how many pairs the sample mode selected for the checker model.
  • checked: how many sampled pairs carry a graded verdict. This is the number that is billed.
  • not_checked: pairs that were not graded.
  • verdicts: a count of each of the four verdicts.

Then comes rows[]. A receipt from a Hyperresearch run also carries a repair block, which records what the citation repair pass fixed before the report shipped and what it left. A standalone verification grades and does not edit, so its receipt has no repair block.

Each row carries pair_index, the sentence verbatim, the citation as n or note_id, the verdict, an evidence span quoted from the source, strong, which records whether the sentence contained a strong marker, and checked, which records whether a checker graded the pair.

evidence is the part worth reading. A verdict without the span it rests on is an opinion. The row gives you the text from the source that the checker used to decide, so you can disagree with it.

The four verdicts, with a worked example each

The verdict vocabulary is supported, partially-supported, unsupported and wrong-source. Every example below is invented to show the shape of each verdict. None of them is a real customer document.

supported. Suppose a report says “Global installed battery storage capacity reached 42 GW in 2024 [7]”, and source 7 is a market report whose text reads “installed capacity stood at 42 GW at the end of 2024”. The number, the unit and the year all appear in the cited source. The checker returns supported with that sentence as the evidence span.

partially-supported. Suppose the report says “Battery storage capacity reached 42 GW in 2024, driven mainly by grid-scale projects in China [7]”, and source 7 confirms the 42 GW figure but says nothing about what drove the growth or about China. Half of the sentence is in the source and half is not. The verdict is partially-supported, and the evidence span shows you the half that is there. This is the most common finding on a competent document, and it is usually fixed by splitting the sentence or moving the citation.

unsupported. Suppose the report says “Battery storage capacity grew 60% year over year in 2024 [7]”, and source 7 gives only an absolute 42 GW figure with no prior-year comparison and no growth rate. Nothing in the cited source establishes the claim. The verdict is unsupported. Because that sentence contains % and grew, it was never eligible for the mechanical auto-pass, so it was always going to be read.

wrong-source. Suppose the report says “Battery storage capacity reached 42 GW in 2024 [7]”, the 42 GW figure really does appear in the corpus, but source 7 is a paper about lithium extraction permitting and the figure came from source 9. The claim is true and the pointer is wrong. The verdict is wrong-source, which is the verdict that catches a renumbering error, a shifted reference list, or a model attributing a number to the nearest citation rather than the correct one.

Under the default ship-gate rules, any unsupported or wrong-source pair fails the gate unless the request sets policy.allow_unsupported. partially-supported is reported and does not fail by itself.

Quote integrity checks the words inside the quotation marks

quote_integrity reports quotes, verified, altered and not_found, and a row per quoted span. Each row carries the quote, the note_id it was checked against, a status of verified, altered or not-found, and for a near-miss the closest span found in the source plus an edit distance.

Every quoted span of six words or more is searched in the cited source body, normalised for whitespace, hyphenation and ellipses, so a line break inside a quotation is not reported as an alteration. A quote that cannot be found verbatim in any source body is a hard block in a Hyperresearch run: a hallucinated quotation cannot ship.

The retraction sweep blocks shipping, not just reporting

retraction_sweep reports dois_checked, retracted, concerns, unavailable, and a row per DOI with doi, a status of ok, retracted, expression-of-concern, corrected or unavailable, the notice date, a notice_url, and noted_in_document.

The retracted status comes from OpenAlex’s is_retracted flag, which the lookup queries first. When OpenAlex has no record, the lookup falls back to Crossref. That fallback reads the wrong field, so it does not currently detect retractions. expression-of-concern and corrected are in the schema but are never produced today. So a swept DOI comes back retracted, ok or unavailable, concerns is zero, and date and notice_url are null. Retraction-aware research documents the gap in full.

noted_in_document is what makes the gate usable. Citing a retracted paper is legitimate when the document says the paper was retracted, which is exactly what a paper about research integrity does. So the gate fails on a retracted citation the document does not itself describe as retracted, and passes when the retraction is acknowledged near the citation.

The independence audit counts voices, not URLs

independence reports sources, clusters, effective_sources, and a row per cluster with cluster_id, a kind of url, body or wire, the root source, and the members list.

Five outlets running the same press release are five URLs and one voice. The audit clusters them by URL identity, by body similarity, and by wire-service origin, then reports an effective source count that treats each cluster as one. A document with 40 sources and an effective count of 31 has nine sources that are repeats of something else. Independence findings are reported and do not fail the gate by default.

What is checked in full and what is sampled

An extractor binds each citation marker to the sentence that carries it, producing citation-sentence pairs. A triage pass then auto-passes any pair whose numbers or a long word-overlap window already appear in the cited source’s extracted claims. Those pairs are free and are counted in auto_passed. The remainder goes to the cite-checker model in batches of up to 25 pairs.

Two sampling rules matter to anyone reading a receipt:

  • Sentences the pipeline marked as strong claims are always read. A quantitative sentence never gets waved through on word overlap.
  • Everything else depends on sample. sample: all reads every pair, and that is the default. sample: strong-markers reads the strong pairs plus a seeded 20% of the rest. A numeric sample value between 0 and 1 reads that share of the non-strong pairs.

This is why pairs, auto_passed, sampled_for_llm, checked and not_checked are on the receipt and not buried. The counts tell you how much of the document was actually read by a checker, so nobody has to guess.

What a receipt does not guarantee

A receipt records checks for review. It is not a guarantee that every claim in the document is true.

Four specific limits follow from that:

  • A supported verdict means the cited source says what the sentence says. It does not mean the source is correct. A receipt cannot tell you that a peer-reviewed paper’s method was flawed.
  • Under a sampled run, pairs outside the sample were not read by a checker. The receipt says how many.
  • Verification is text-only. Figures, tables published as images, and audio are out of scope, so a claim resting on a chart is not verified by the quote or cite-check tables.
  • The extractor reads numbered [N] and wikilink [[note-id]] citations. Author-year and footnote styles are not supported today.

Reading a receipt in the console

A finished run’s detail page has a Receipt tab. It renders the ship-gate checklist first, then the four tables, and it is built to print, because the common use is attaching it to something. The report tab uses the same data from the other direction: citation markers in the report carry their verdict on hover and open the source note on click, and sentences the patcher edited are marked.

Reading a receipt from the API

For a standalone verification, poll GET /v1/verify/{id} until status is terminal, then read GET /v1/verify/{id}/result. That response carries id, status, receipt_url (a signed reference to the durable copy in storage, valid about an hour) and receipt, which is the full object inline.

For a research run, the receipt arrives as the verification member of the run result, alongside report.md, sources[], claims[], telemetry and the run’s artifacts.

Both are the same schema. A receipt from a $49 Deep run and a receipt from a document you pasted into the verification API can be read by the same code.

Next: the citation verification API for running these checks on your own documents, and pricing for what a run and a checked citation cost.

By Jordan Gibbs · Updated 2026-09-23