Guides

Claude Code runs deep research when you give it a harness

Claude Code can run a research job that lasts an hour or more and reads dozens of sources, but one prompt will not get you there. It needs a harness: procedure files it loads one step at a time, subagents that do the reading in their own context, tool limits on the subagents that edit the report, and notes kept on disk. The open-source hyperresearch CLI is one such harness. It is MIT licensed, installs with pip, and has about 3.6k stars on GitHub as of September 2026: https://github.com/jordan-gibbs/hyperresearch. This guide sets it up and shows how a run is built, so you can use it or copy the design.

Install and run

You need Python 3.11 to 3.14 and Claude Code. In the folder where you want the research to live:

Terminal window
pip install hyperresearch
hyperresearch install

That writes the skills and subagents into the project. If Claude Code was already open, restart it in the same folder. Then, inside Claude Code:

/hyperresearch What changed in municipal fibre financing after 2023?

You can also install it as a Claude Code plugin with /plugin marketplace add jordan-gibbs/hyperresearch and then /plugin install hyperresearch@hyperresearch. The plugin still needs the pip package.

The default fetcher is plain HTTP. The pipeline is tuned for crawl4ai, a headless browser, which is an optional extra:

Terminal window
pip install "hyperresearch[crawl4ai]"
.hyperresearch/config.toml
[web]
provider = "crawl4ai"
[scholar]
contact_email = "you@example.com"

Set contact_email too. Unpaywall requires one, and without it open-access recovery only checks Europe PMC.

Skills load one step at a time

Claude Code keeps a skill’s description in context and loads the full text only when the skill is invoked (see https://code.claude.com/docs/en/skills). The /hyperresearch entry skill uses this. Its own description calls it “a ROUTER”: it sets the run up, then invokes one step skill per phase, from hyperresearch-1-decompose to hyperresearch-16-readability-audit.

The skill file says why. An earlier version was one long skill, and after context compaction “the orchestrator forgot the procedure, wrote a single draft, and produced a flat-scoring report.” Loading each step when it runs puts the procedure in context at the moment it is needed.

Step 1 sorts the question into light or full. A light run does five steps. A full run does all sixteen plus a citation check.

The run writes down what it must not forget

Before step 1, the entry skill creates the vault, makes a unique run tag, and writes a manifest at research/runs/<tag>/run.json. The orchestrator records every step there. That is why a crashed run can resume: hyperresearch run resume -j returns the exact next step.

The question is saved word for word to research/runs/<tag>/query.md. Every step and subagent re-reads it from that file.

Subagents do the reading, and two cannot rewrite

A Claude Code subagent has its own context window, and its definition can set its model and the tools it may use (see https://code.claude.com/docs/en/sub-agents). hyperresearch ships sixteen of them. Fetchers run 8 to 12 at a time, so long pages stay out of the orchestrator’s context. Four critics attack the draft in parallel. A cite-checker tests a sample of citation and sentence pairs before the report ships.

The patcher and the polish auditor may only use Read and Edit. They cannot write a new file, so they can apply the critics’ findings only as small edits. This matters because the cheapest way to satisfy a critic is to delete the passage it objected to, evidence and all. Taking away the Write tool takes away that option. If you build your own harness, copy this part.

Check the run before you trust the report

Terminal window
hyperresearch run status -j # current step, spend, escalation queue
hyperresearch run report -j # wall time, spend and sources per step
hyperresearch run verify <tag> -j # ship gate

The ship gate checks headings, length, citation density and whether the cite-check findings were resolved. A quote that does not appear word for word in a vault note blocks the report. So does citing a retracted source without saying so. Numbers that cannot be traced to evidence are flagged. The README says these checks “cannot guarantee factual accuracy.”

Every source lands in research/notes/ as a Markdown file with YAML frontmatter. The SQLite index next to it is a cache, and hyperresearch sync rebuilds it. hyperresearch search "<query>" -j searches the notes, and the next run on the same topic checks the vault before it goes to the web.

Limits of running it yourself

The README gives typical times of about 30 to 40 minutes for a light run and 1.5 to 2.5 hours for a full run. Those are estimates, not benchmarks. The steps run one after another in one Claude Code session, so that session stays open for the whole run.

Model usage is billed to however your Claude Code is set up. hyperresearch run init <tag> --budget 50 caps estimated spend and stops the run at the ceiling.

When a fetch hits a login wall or a bot wall, the URL goes into an escalation queue instead of failing the run. With the Claude in Chrome extension, a browser-fetcher subagent can work through the queue in your own Chrome. CAPTCHAs, two-factor prompts and logins are never solved automatically. They come back to you.

The hosted version

If you would rather not run it yourself, Hyperresearch runs the same pipeline as a hosted service. You start a run with a REST call to https://api.hyperresearch.ai/v1 or from the remote MCP server, and it does not tie up your machine. It is paid per run, Light $9 and Deep $49 (pricing). The CLI page compares the two line by line. The skill files in the repo are the specification both follow.

By Jordan Gibbs · Updated 2026-09-27