Product

DRACO: how we tested

We wanted a fair answer to one question: on real research tasks, graded on accuracy, how do Hyperresearch’s reports compare with other research agents? This page explains what we ran, how it was graded, the per-task scores, and the limits of the result.

What DRACO is

DRACO is a 2026 deep research benchmark from Perplexity and Harvard. It has 100 real research tasks across 10 domains, each with grading criteria written by experts (26 in all). Factual accuracy is about half of the score, and wrong claims cost points, so a long report full of errors does not score well.

What we ran

Every system got the same 10 tasks, one from each of DRACO’s 10 domains:

  • Finance: an analysis of institutional Bitcoin flows.
  • Shopping: choosing between three multi-axis machining centres for titanium aerospace parts.
  • Academic: the mechanisms of cellular senescence.
  • Technology: data-loss prevention for ChatGPT Enterprise.
  • General knowledge: women’s labour force participation in four countries since 1970.
  • UX design: music software input for musicians with hand tremor.
  • Personalized assistant: the return on three routes from backend engineer to machine-learning engineer.
  • Medicine: monitoring, insurance authorization and pacing thresholds for a patient with a worsening heart block.
  • Needle in a haystack: finding one specific restaurant promotion and the details around it.
  • Law: a buyer’s rights after accepting and using a wrong paper delivery.

We also ran one extra medicine task, a patient with Crohn’s disease deciding whether to go to the emergency room. It is shown separately and is not in the 10-task mean.

We ran the tasks in rounds: 3 first, then 3 more, then the last 4 and the extra task.

Each system received the task prompt exactly as written. If a system asked a clarifying question, it was told: “No preference. Please proceed with your best judgement.”

How it was graded

We used the paper’s per-criterion grader with GPT-5.2 as the judge, at the paper’s settings (reasoning off, temperature 0). The paper’s main judge, Gemini-3-Pro, is no longer available; GPT-5.2 is the paper’s own second judge.

Our reports and the Claude, ChatGPT and Gemini reports were graded twice and the two grades averaged, as were most Perplexity reports. Valyu, Parallel and You.com were graded once, from the reports they published for DRACO.

Our Finance, Academic and Law reports were produced by our current Deep writer from the research each task’s run collected.

DRACO, 10 tasks, same judge

System Tasks Mean Medicine, patient decision (extra task)
Hyperresearch Deep 10 of 10 77.7 89.5
Claude Research (Opus) 10 of 10 73.0 90.2
Parallel Ultra (published reports) 10 of 10 61.5 85.0
You.com Research (published reports) 8 of 10 60.1 76.7
Perplexity Deep Research 8 of 10 57.8 84.2
Valyu DeepResearch Heavy (published reports) 8 of 10 57.2 80.5
ChatGPT Deep research † 10 of 10 55.2 39.1
Gemini Deep Research ‡ 10 of 10 54.7 46.6

Hyperresearch is top of the board, 4.7 ahead of Claude Research (Opus). Both are more than 11 points ahead of every other system.

Scores by task

System Finance Shopping Academic Technology General knowledge UX design Pers. assistant Medicine Needle Law
Hyperresearch Deep 70.8 80.2 91.2 90.1 65.9 74.2 64.1 72.2 77.6 90.6
Claude Research (Opus) 68.5 80.9 91.8 86.6 69.0 80.6 54.6 37.1 64.2 96.9
Parallel Ultra 45.8 74.6 75.5 69.4 50.4 50.5 62.6 29.4 71.6 85.0
You.com Research 14.2 83.7 78.2 51.7 67.2 44.3 77.3 64.4
Perplexity Deep Research 27.1 62.3 89.0 66.9 28.0 22.2 79.6 87.5
Valyu DeepResearch Heavy 54.2 87.8 71.5 72.6 26.4 0.0 81.1 63.7
ChatGPT Deep research † 27.8 72.4 67.4 69.4 41.5 33.8 37.0 52.6 68.7 81.3
Gemini Deep Research ‡ 38.7 54.0 78.4 72.6 43.4 21.6 45.5 35.1 78.6 79.4

Empty cells:

  • You.com, Shopping and Needle: its published report for these tasks is a copy of Valyu’s.
  • Valyu, Shopping and General knowledge; Perplexity, General knowledge and UX design: not graded in this round.

Claude, ChatGPT and Gemini were run in their consumer apps with no custom instructions and graded twice. Parallel, Valyu and You.com were graded once, from the reports they published for DRACO. Most Perplexity reports were graded twice, a few once. Gemini’s later tasks ran on 3.8 Flash, its top Deep Research model on October 5.

† ChatGPT’s reports were exported from the app without source links, so its citation scores reflect that.

‡ Gemini’s Finance report came from the Gemini API; the rest came from the Gemini app and were copied out with their source titles but without links, so their citation scores reflect that.

Sources

The score says how good a report is. These counts show what it rests on. Each is an average per run on the first 3 tasks (Finance, Academic and Technology), from our runs on those tasks.

Hyperresearch Deep Perplexity Deep Research ChatGPT Deep research Gemini Deep Research
Distinct sources cited per report 73 29 not exported 29 (Finance)
Cited figures checked against the source, per report 96 none none none
Sources kept for your next run 223 none none none

How we counted. For Hyperresearch, “cited” is the number of distinct sources the report’s citations resolve to, from each run’s verification receipt; “checked” is the number of cited figures the claim check read against the full text of the source they cite; “kept” is the number of sources the run read, every one of which stays in your vault for the next run in the same project. For Perplexity, “cited” is the distinct links in the citation list its API returned; for Gemini, the entries in its numbered source list. ChatGPT’s reports were exported without source links, so there was nothing to count. “None” means the product has no check of each cited claim against its source and no library of sources that a later run searches, as of October 2026. We have not counted sources in the Valyu, Parallel and You.com reports.

Limits

  • These are 10 of DRACO’s 100 tasks, plus one extra.
  • Two grades of the same report can differ by up to about 9 points.
  • This is our measurement, not an official leaderboard.
  • These tasks are easier than DRACO’s average, so every system scores higher here than in the paper’s full-set results. The two sets of numbers cannot be compared. Model scores that labs report from their own DRACO setups use different judges and harnesses and are not comparable either.

Date and versions

October 2026. Every system on this page changes over time, ours included. We will re-run the tasks and update this page.

Updated 2026-10-06