Alternatives
Migrating off the OpenAI deep research API
OpenAI shut down o3-deep-research and o4-mini-deep-research on 2026-07-23. Its deprecations page names gpt-5.6-sol as the substitute for both. That is a general model, not a research model, so moving to it means you build the research loop yourself. The other options are a task API from another vendor, or a research run from Hyperresearch at $9 or $49 a report. Which is right depends on what you need back from each call.
OpenAI’s deep research guide at developers.openai.com/api/docs/guides/deep-research still describes both models as current, with no deprecation notice, as of 2026-09-27. Plan from the deprecations page.
What the old call did
One Responses API request with a research model and at least one data source:
POST https://api.openai.com/v1/responses{ "model": "o3-deep-research", "input": "<question>", "background": true, "tools": [{ "type": "web_search_preview" }]}The model searched, read and wrote a report. Citations came back as annotations on the output text, each with url, title, start_index and end_index. You could add file_search with vector_store_ids to read your own documents, code_interpreter to compute, or a remote MCP server. Runs could take tens of minutes, so the guide said to use background: true and poll the response id.
Staying on OpenAI
This is the smallest change. Swap the model for gpt-5.6-sol, keep the Responses API, and attach the web search tool. The model supports web search, file search, code interpreter and MCP, and background mode with polling on GET /v1/responses/{id} is still there.
Published prices on 2026-09-27:
- Tokens: $4.00 per million in, $0.40 per million cached in, $20.00 per million out. Long-context prompts bill at $8.00 in and $30.00 out. OpenAI marks this as promotional pricing “available at least through November 21, 2026”.
- Web search: $10.00 per 1,000 calls, plus the search content billed as tokens at the model’s rates.
- File search: $2.50 per 1,000 calls.
You pay for what your loop uses, so the cost depends on how many searches it runs and how much text it reads. As an estimate, a loop that makes 50 searches, reads 1 million input tokens and writes 50,000 output tokens costs:
50 searches x $10.00 / 1,000 = $0.501,000,000 in x $4.00 / 1M = $4.0050,000 out x $20.00 / 1M = $1.00 ----- $5.50Those token counts are assumptions for illustration, not measurements. Cached input at $0.40 lowers the input line if your loop re-reads the same context.
What you give up is the research behaviour itself. You now decide how to break the question down, how many searches to run and when to stop. The model returns what it writes, and checking that each citation supports its sentence is your job.
Moving to another vendor’s research API
Some vendors sell a research task as one call. Parallel’s Task API is priced per 1,000 requests; its Ultra processor, labelled “Extensive deep research”, is $300 per 1,000, or $0.30 a task, and Ultra8x is $2,400 per 1,000. Perplexity’s sonar-deep-research bills per token: $2 per million in, $8 out, $2 for citation tokens, $3 for reasoning tokens, and $5 per 1,000 searches. Both figures come from the vendors’ pricing pages on 2026-09-27.
These keep the “one call, one answer” shape of the old API at a per-call price well under a dollar.
Moving to Hyperresearch
A Hyperresearch run takes one question and returns a long report, the sources behind it and the claims it relies on; a Deep run adds a verification receipt. The receipt lists the sampled citations, each checked against the source it cites, with a verdict of supported, partially supported, unsupported or wrong source. Retracted papers and quotes that do not appear in their source stop a report from shipping.
Two sizes, flat per run:
- Light, $9: about 20 to 50 sources and 2,500 to 4,000 words, in about 30 to 40 minutes.
- Deep, $49: about 150 to 200 sources, split into 3 to 5 chapters, in 3 to 5 hours.
Source counts and times are the targets the pipeline gates on, not guarantees. The price is fixed when the run is created, not metered by tokens.
On ten DRACO tasks, Hyperresearch scored 77.7
The retired API models can no longer be tested, so the nearest OpenAI product we could grade is Deep research in the ChatGPT app. We ran it, Hyperresearch Deep, Perplexity’s sonar-deep-research API and other research agents on the same tasks from DRACO, a deep research benchmark from Perplexity and Harvard, in October 2026, and graded each with the same judge, GPT-5.2 at the paper’s settings, mostly as the mean of two grades. Valyu DeepResearch Heavy, a research task API, was graded once from the reports Valyu published for DRACO.
| System | DRACO tasks | Mean score |
|---|---|---|
| Hyperresearch Deep | 10 of 10 | 77.7 |
| Perplexity Deep Research (API) | 8 of 10 | 57.8 |
| Valyu DeepResearch Heavy | 8 of 10 | 57.2 |
| ChatGPT Deep research (app) | 10 of 10 | 55.2 |
Against ChatGPT the widest gaps are on the Finance task, 70.8 against 27.8, and the UX design task, 74.2 against 33.8. ChatGPT came closest on the Shopping task, 72.4 against 80.2. Valyu and Perplexity scored higher than us on the Needle in a haystack task. ChatGPT’s reports were exported from the app without source links, so its citation scores reflect that. Perplexity’s and Valyu’s means cover the 8 tasks with a graded report. These are 10 of DRACO’s 100 tasks, two grades of the same report can differ by up to about 9 points, and this is our measurement, not an official leaderboard. How we tested has every task score and the other systems we graded, including Claude Research (Opus) at 73.0, or see a full Deep report.
The shapes side by side
OpenAI with gpt-5.6-sol |
Research task API | Hyperresearch | |
|---|---|---|---|
| You pay for | Tokens plus search calls | Per task, or tokens plus searches | Per run, $9 or $49 |
| Research loop | Yours to build | The vendor’s | The pipeline’s |
| Citations | annotations with character offsets |
Varies by vendor | Every kept source in your vault, with a verification receipt on Deep |
| Citation checking | Yours to build | Varies by vendor | Built in |
| Time per call | Depends on your loop | Varies by processor | About 30 to 40 min Light, 3 to 5 h Deep |
A $0.30 task and a $49 run are different purchases. Decide which one you need before comparing prices.
The call, field by field
curl -X POST https://api.hyperresearch.ai/v1/runs \ -H "Authorization: Bearer <YOUR_API_KEY>" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: migrate-0001" \ -d '{ "query": "<question>", "tier": "premier", "levers": { "register": "analyze", "purpose": "decide" }, "webhook_url": "https://example.com/hooks/hyperresearch" }'| Old | New | Note |
|---|---|---|
input |
query |
Up to 4,000 characters |
model |
tier |
light or premier. premier is Deep on the wire. Omitted, it is light |
background: true |
nothing | Every run is asynchronous. The POST answers 202 |
tools: [web_search_preview] |
nothing | Searching and fetching are part of the run |
file_search with vector_store_ids |
your workspace vault, and project_id |
Runs read the vault before the web. A project_id scopes that reuse to one project |
| polling the response id | GET /v1/runs/{id}, or webhook_url |
Use either |
| tone and structure in the prompt | levers |
register, purpose, reader_expertise, report_language, response_format, required_section_headings, domain_notes |
| none | Idempotency-Key |
Same key and body returns the original run. Same key with a different body is a 409 |
The 202 gives you the run id and the price before any work starts:
{ "id": "run_01J9ZQ7V6H0000000000000000", "status": "queued", "tier": "premier", "price_cents": 4900, "events_url": "...", "created_at": "2026-09-18T14:02:11Z"}Waiting for the run
Poll GET /v1/runs/{id} for the status, the current step, steps_done out of steps_total, sources kept so far, and on a failed run a failure_reason. A run can also stop at blocked and wait for a decision, for example a source it could not read without your help. blocked_code tells you which gate it stopped on.
To skip polling, set webhook_url. Events include run.started, run.step, run.blocked, run.done, run.failed and verify.done.
When the run is done, GET /v1/runs/{id}/result returns report_markdown, sources, skipped (every page tried and not kept, with a reason), claims, verification, escalations and artifacts. Artifact links are signed URLs valid for one hour.
If an agent is calling rather than a backend, the same start, check and fetch pattern is available over MCP at https://mcp.hyperresearch.ai/mcp. Long-running MCP tools explains why it takes three calls.
When OpenAI or a task API is the better choice
Stay on OpenAI if you already have a research loop, or want full control of it, and your costs per call land in the single dollars or below. The estimate above came to $5.50 for a fairly heavy loop. You also keep token-level control, character-offset citations and one vendor.
Use a task API if you need an answer while a user waits, or if you run thousands of lookups a day. A Hyperresearch run takes about half an hour or more and costs at least $9. No amount of checking makes that the right purchase for a quick lookup.
Hyperresearch fits when the report is the thing you deliver and someone will be held to what it says. There, the difference between a citation that exists and one that supports its sentence is worth paying for. Pricing has the full numbers.
By Jordan Gibbs · Updated 2026-10-06
