Guides
Cost per verified citation is the unit this market lacks
Cost per verified citation is the price of one delivered research task divided by the number of citations in that deliverable whose cited source the provider actually read and compared against the sentence citing it. It exists because this market has no shared unit. Vendors publish per 1,000 tasks, per 1,000 searches, per million tokens and per page, and none of those tell a buyer what a finished, checkable report costs. The unit is also flattering to Hyperresearch, the AI deep research API, for a reason stated below, so read the method and compute cost per task alongside it.
The definition, and what it excludes
A citation counts as verified when three things are true of one delivered task:
- The provider retrieved the body of the cited source, not its metadata record.
- The provider compared that body against the specific sentence carrying the citation.
- The result is a per-citation verdict in the deliverable, readable by the buyer without asking.
Five things do not count, and each one is a real check that some tool performs:
- Existence and metadata. A DOI that resolves and a title that matches Crossref tells you the reference is real. A retracted paper resolves perfectly.
- Retraction status. A register lookup is a different question from support.
- Quote integrity. Confirming that a quoted span appears verbatim in the source checks the words, not the claim the sentence makes with them.
- Unsampled citations. If a provider checks a sample, only the sample counts.
- Unresolvable markers. A citation marker with no source behind it is a finding, not a verified citation.
Two variants follow, and both should be published together:
- Strict. Only citations read by a model against the source text. In a Hyperresearch verification receipt this is the
checkedcounter. - Inclusive. Strict plus citations that cleared a mechanical comparison against the source’s extracted claims, which is the
auto_passedcounter.
Publish strict. Inclusive is the honest upper bound, and a provider that publishes only the inclusive number is doing something with it.
The denominator is a counter, not an estimate
The Hyperresearch verification API bills by the citation-sentence pair that has to be read, and pairs that clear the mechanical triage are free. The billed event and the denominator are therefore the same event, so there is nothing to estimate: the receipt reports pairs_total and pairs_llm, and the amount charged is on the same response. Quote integrity, the retraction sweep and the independence audit are included and do not enter the denominator. The citation verification API has the request shape.
For a full research run, the numerator is known and the denominator is not
Flat per-run prices and source targets are published. Citations per report are not.
| Run | Price | Sources, target | Price per source in the corpus |
|---|---|---|---|
| Light | $9 | 20 to 50 | $0.18 to $0.45 |
| Deep | $49 | 150 to 200 | $0.25 to $0.33 |
Deep’s source counts are the targets the width sweep gates on and Light’s are what production Light runs keep today; neither is a guarantee, so the right column is a band and not a price. More importantly, a source is not a citation. One source can carry ten citations and another can carry none. Until pairs, auto_passed and checked are measured across a real sample of runs at each size, there is no honest per-size cost per verified citation to publish, and this page does not publish one.
What the receipt already gives any buyer is the ability to compute it themselves after one run: divide the run price by the receipt’s checked count.
The metric is undefined for the cheapest run
A Light run checks its cited claims against the notes they cite and corrects what does not hold before it ships, but it carries no verification receipt, so no per-citation verdict reaches the buyer. By the definition above it delivers zero verified citations, so its cost per verified citation is undefined no matter what it costs. The cheapest thing on the price list scores worst on the metric its vendor invented. That is the correct behaviour for an honest unit, and it is the first thing to check when someone shows you one.
Per-call APIs score undefined too, and that is the trap
The same division by zero applies to every research API that does not verify. Checked against each vendor’s own pricing page on 2026-09-17:
- Parallel Task API. Priced per 1,000 requests across nine processor tiers, from $5 to $2,400 per 1,000, so $0.005 to $2.40 per task. No citation verification is offered, so the denominator is zero.
- Exa. The Agent product is priced per request at fixed effort levels from $0.012 to $1.00, or metered at $0.10 per Agent Compute Unit plus $0.005 per search tool call. Search and contents are separate line items. Denominator zero.
- Perplexity.
sonar-deep-researchis billed per token: $2 per million input, $8 per million output, $2 per million citation tokens, $3 per million reasoning tokens, plus $5 per 1,000 searches. There is no per-report price on the page, so cost per task requires running it and reading the usage back. Denominator zero. - Tavily. Research is 15 to 250 credits per request at
model=proand 4 to 110 atmodel=mini, with pay as you go at $0.008 per credit, so $0.12 to $2.00 per pro request. Denominator zero. - OpenAI. No deep-research model appears on the pricing page today. The deprecations page records
o3-deep-researchando4-mini-deep-researchshut down on 23 July 2026, announced 2026-04-22, withgpt-5.6-solnamed as the replacement.
Presenting that as an infinite cost per verified citation would be dishonest, and anyone who does it is selling you the metric rather than the product. The correct reading is narrower: these are priced per call and they answer a different question. A Parallel task at half a cent and a Hyperresearch Deep run at $49 differ by four orders of magnitude because one returns an answer and the other returns a reviewed document with a ship gate. If the per-call answer is what your product needs, the per-call price is the number that matters and this unit is irrelevant to you.
So publish two columns, always: cost per task, and cost per verified citation. The first is where per-call APIs win by a wide margin. The second is where the difference in what you receive becomes visible.
The method, so anyone can run it against us
Reproducible in a day, by a competitor, with no cooperation from any vendor.
- Fix the task set. Twenty questions, written down before any run, spanning at least four domains, with no question chosen because a system handles it well.
- Fix the deliverable. One document per question per provider, obtained through each provider’s documented API at its documented settings. Record the settings.
- Record the invoice, not the estimate. Numerator is what the provider billed for that task, including metered add-ons and search charges. For token-billed providers, read usage from the response.
- Count the denominator by hand on a sample. For each provider, take ten citations per document at random. For each, open the cited source, find the passage the provider claims supports the sentence, and record whether the provider itself stated a per-citation verdict. Only citations with a stated, checkable verdict count.
- Extrapolate once, and say so. Scale the hand-counted rate to the document’s citation count. Publish the sample size beside every derived number.
- Publish both variants and both columns. Strict and inclusive, cost per verified citation and cost per task.
- Publish the failures. Tasks that did not complete, documents with no citations, and providers whose deliverable made step 4 impossible.
Step 4 is the step that cannot be skipped and the step every vendor comparison skips.
What a verified citation is not
A verified citation is a checked citation, not a correct one. A supported verdict means the cited source says what the sentence says. It does not mean the source is right, and it does not mean the checker was right.
Hyperresearch has not measured its own cite-checker’s false-positive rate on a defined sample. Until that number exists, every denominator on this page is a count of checks performed, not of checks known to be correct, and the metric inherits whatever the error rate turns out to be. Treating a high verified-citation count as a quality score is exactly the mistake this unit was meant to stop.
Next: what a verification receipt contains, which is where the counters in the denominator come from, and verifying citations in AI-generated text for what each verdict means.
By Jordan Gibbs · Updated 2026-09-23
