Guides

How to run a job longer than an MCP tool call

An MCP tool call is a request and a reply inside a live client session, so it inherits that session’s timeout and that client’s result size limit. If your tool does work measured in tens of minutes, or returns a document measured in hundreds of kilobytes, it cannot be one call. The fix is to split it: one tool starts the work and returns a handle, a second tool reports progress, a third returns the result. This page gives the numbers, the three designs and the reasons to pick one, the exact contract Hyperresearch uses, and a recipe you can copy into your own server.

The limits are real, and mostly undocumented

Two separate limits bite, and they are documented unevenly.

Time. Claude Code’s per-tool-call timeout is MCP_TOOL_TIMEOUT in milliseconds, and its documented default is “inherited from MCP_TIMEOUT”, which itself defaults to 30000, so an unconfigured Claude Code aborts a tool call after 30 seconds. For claude.ai and Claude Desktop the figure usually quoted is about 240 seconds. That number is not stated on any Anthropic-owned page. Searching docs.claude.com, code.claude.com, platform.claude.com and support.claude.com on 2026-09-17 found it only in GitHub issues against anthropics/claude-code and in third-party posts, and the Messages API MCP connector page contains no per-call timeout at all. It is an observed figure, not a documented one. Treat 240 seconds as a ceiling you should stay far below, not a contract you can rely on.

Size. Here Anthropic’s own documentation disagrees with itself. The MCP page says Claude Code warns when a tool’s output exceeds 10,000 tokens and that “the default maximum allowed MCP output is 25,000 tokens”, configurable through MAX_MCP_OUTPUT_TOKENS. The environment-variable reference for the same variable says the default is 300000 and that “when an MCP tool returns more tokens than this limit, Claude Code truncates the output and reports the truncation in the tool result”, adding that “exceeding the token limit can cause model requests to fail.” Both pages were read on 2026-09-17, and the conflict is Anthropic’s, not a gap in this page. For claude.ai and Claude Desktop the commonly quoted cap is around 150,000 characters, which is again an observed figure with no primary source.

Two things follow whatever the exact numbers are. A server author cannot know the caller’s limit, and the caller’s limit can be lower than the author expects by a factor of ten. Design for the small number.

One documented lever exists. A server can raise its own per-tool threshold by setting _meta["anthropic/maxResultSizeChars"] on the tool’s tools/list entry, and Claude Code will honour it “up to a hard ceiling of 500,000 characters”. Above the threshold Claude Code “saves it to a file and replaces it in the conversation with a message that names the file path”. That helps a large result. It does nothing for a long one.

Three designs, and why two of them fail

Block. The tool does the work and returns the answer. This is correct and should be your first choice for anything that finishes in a few seconds. It fails the moment the work outlasts the smallest timeout in your caller population, and it fails silently from the model’s point of view: the call comes back as an error with no way to recover the work that was in flight.

Stream. The tool sends notifications/progress against the progressToken the client put in _meta, and returns when done. Streaming solves the user experience problem, because the client can show movement, and on some clients a live stream postpones an idle timeout. It does not solve the deadline. The call is still one call, the total is still capped, and clients are free to ignore progress notifications entirely. Stream as a courtesy, never as your durability story.

Start and poll. One tool starts the work and returns a handle immediately. The work runs somewhere that is not the request. Other tools read status and result against the handle. Every call is short, every result is bounded, and the work outlives the session, the client and the network.

Hyperresearch uses start and poll because the unit of work makes the choice for us. A Light run takes about 30 to 40 minutes and a Deep run targets 3 to 5 hours. Neither is a guarantee, and even the faster one exceeds the 240-second figure, let alone a 30-second default. A finished report is also far larger than 25,000 tokens. There is no version of a blocking tool that works here.

The contract

Three tools, three jobs.

start_run returns a handle and a price. Arguments are query, tier, optional profile, levers, budget and idempotency_key. It returns immediately:

{
"id": "run_01J9ZQ7V6H0000000000000000",
"status": "queued",
"tier": "premier",
"price_cents": 4900,
"events_url": "...",
"created_at": "2026-09-18T14:02:11Z"
}

The price is fixed at start, before any work happens, and returned in the same response as the handle. A tool that spends money should tell the model what it just spent in the reply that commits the spend, not in a later call the model might not make. start_run requires the runs:write scope, which is a separate browser grant from the read-only default.

run_status returns a step and progress counters. Given run_id, it returns the status, the current step by id and name, steps_done against steps_total, sources_fetched, notes_written, candidates_seen, spend so far, and queued and resolved escalation counts. A terminal run carries finished_at, duration_ms and, when it ended badly, one customer-safe failure_reason plus a stable blocked_code to branch on. Keep this response small. It is the call that gets made forty times.

run_result returns the report and the receipt, in parts. Given run_id and a part of report, sources, verification, claims, artifacts, escalations or telemetry. The report part returns the first 20,000 characters, a truncated flag, and a signed URL valid for one hour for the whole thing. On a Deep run, the verification part returns the receipt: every citation checked against the source it cites with a verdict of supported, partially supported, unsupported or wrong source, plus the four counters that say how the run’s own checks went, passed, warning, failed and skipped.

Parts matter more than they look. A caller that wants to know whether a report is safe to quote needs the receipt and not the prose, and a caller writing a summary needs the prose and not the telemetry. Splitting the result by part is what keeps every individual response inside a limit you cannot see.

Idempotency is the price of making the start call retryable

A tool call that charges money and might time out is a double-billing bug waiting for a slow network. start_run takes an idempotency_key of 1 to 255 characters; the REST equivalent takes it as an Idempotency-Key header on POST /v1/runs. Replaying the same key with the same body returns the original run instead of starting a second one. Replaying it with a different body is a 409, because that is a client bug and not a retry.

Generate the key on the client, before the first attempt, and reuse it for every retry of that attempt. A key generated inside the retry loop is not an idempotency key.

When the client disconnects, nothing happens to the run

The run is not the tool call. It executes in the orchestrator as a durable workflow and it is owned by the workspace, not by the MCP session. Closing the client, losing the network, revoking the grant and reconnecting from a different machine all leave the run alone. Come back later, call run_status with the id, and the state is there. A run you actually want to stop is stopped explicitly with control_run, which also pauses and resumes.

For clients that can receive server-initiated messages there is a push path. resources/subscribe on hr://runs/{id}/report produces a notifications/resources/updated when the run finishes. Notifications are buffered per session and replayable: a GET /mcp with Last-Event-ID replays what the client missed and then closes, so a reconnecting client catches up without the server holding an idle connection open. notifications/cancelled aborts an in-flight tool call, which cancels the read and not the run.

Push is an optimisation on top of polling. Polling is the contract, because a client that supports no notifications at all must still work.

The recipe, for any MCP server

  1. Measure your p99, not your median. If p99 is over 20 seconds, assume it will not fit a tool call. The smallest documented default in the field is 30 seconds.
  2. Name the three tools after the job, not after the pattern. start_export, export_status, export_result beats create, poll, get. Tool selection is semantic over your metadata, so the name and description are the whole interface the model sees.
  3. Make the start tool return the handle, the terminal states it can reach, and the cost if there is one. A model that knows the states can branch; one that gets a bare id will guess.
  4. Say the pattern in the tool description, in an imperative sentence. “Returns immediately with a job id. Poll export_status until status is done, then read export_result.” Client models follow instructions in tool descriptions, and this is the cheapest place to put the one instruction that makes the sequence work.
  5. Take an idempotency key on the start tool and honour it. Return the original on a replay of the same body, and an error on a replay of a different one.
  6. Keep the status response tiny and monotonic. One status string from a closed set, a step, a done-of-total pair, and nothing that grows. Do not return partial output from the status tool.
  7. Split the result by part, and hand back a signed URL for anything large with the head of the content inline. Set _meta["anthropic/maxResultSizeChars"] if a part genuinely needs more room, remembering the 500,000-character ceiling.
  8. Send notifications/progress when the client supplies a progressToken, and treat it as decoration. Never let correctness depend on a notification arriving.
  9. Run the work outside the request, in something durable, and key it to the tenant rather than the session. If your job dies when the client disconnects, none of the rest of this helps.
  10. Honour notifications/cancelled by aborting the read, and give the caller a separate, explicit tool to cancel the job. Those are two different intentions and collapsing them will cancel expensive work by accident.

The pattern is not specific to research. It is what any MCP tool needs when the work is a job rather than a lookup: a render, a migration, a large export, a test suite, a build. Copy the shape and put your own nouns in it.

Hyperresearch’s implementation is at /mcp/claude for Claude clients and /mcp/cursor for Cursor, and the same three-step contract is available over REST at https://api.hyperresearch.ai/v1 with webhooks instead of polling.

By Jordan Gibbs · Updated 2026-09-23