When your MCP server returns an error, the first reader is no longer a human engineer scanning a log; it is an agent deciding whether to retry, fix its arguments, or stop. A researcher-run experiment posted to arXiv in late September 2026 measured recovery swinging from 6% to 88% based solely on what the error step told the agent to do. This article turns that evidence into a shippable error-design checklist and a failure-injection protocol you can run against your own server before release.
Who actually reads your MCP errors now
Most tool-server authors write error strings for the audience they know: a developer at a terminal. That produces two familiar styles, stack-trace prose (“Failed to fetch: upstream returned a rate limit after repeated attempts in handler fetchWithRetry”) and terse codes (“E_RATE_LIMIT”). Both assume a reader who can infer context, look up documentation, and choose a remedy.
The MCP specification frames tool-execution errors as feedback that lets the model correct itself and retry, and Anthropic’s guide for tool authors recommends stating specific, actionable improvements. But as the error-message study points out, neither source says who can act on the improvement, nor reports how much a suggested step actually helps an agent recover. The spec gives you the channel; it does not tell you what to say into it.
Meanwhile, most production MCP servers are thin. An analysis of 116 official servers cited in that paper found 88.6% backed fully or partly by REST APIs and 92% implementing tools as bare API wrappers. In practice the error your agent sees is likely a passthrough of an HTTP error written for REST clients, with no thought about what an LLM does next.
One second-order consequence is worth naming before the checklist. Every unclear error you emit shifts the cost of recovery onto every agent client integrating with your server, multiplied by each retry loop. This is the same ownership question that comes up in deciding who owns an agent’s tool layer: whoever frames the failure determines what the agent does about it.
The measured stakes: recovery swings on wording alone
The new paper, “MCP Error Messages Written for Developers Hurt the Most Capable Agents Most” (arXiv:2609.35381), reports two experiments where only the error step’s content changed:
- Rate limit. Replacing the generic instruction “Wait before retrying.” with the named call to repeat raised recovery from 6% to 88%, at about one extra tool call.
- Expired credentials. An error step containing a terminal command left 45% of tasks recovered. Naming the login tool in its place raised recovery to 84%, and deleting the step entirely with a one-sentence prompt raised it to 82%.
Two details in those numbers deserve attention. First, the terminal-command variant is exactly what a developer-minded author writes: technically correct instructions for a human with a shell. It was the lowest of the variants the paper reports for expired credentials. Second, on credentials, saying less (82%) nearly matched saying the right thing (84%) in the same experiment, while saying the developer-oriented thing (45%) was far worse. The failure was not verbosity; it was guidance the agent could not execute.
The paper also reports, citing model guidance from OpenAI and Anthropic, that GPT-6 Astra and Claude Fable 5 weigh written instructions more heavily than their predecessors, so guidance written for earlier models can over-constrain newer ones. That behavior will shift with each model release, so treat it as a reason to re-run your own tests per model generation, not as a fixed design rule.
These recovery deltas are author-reported and unreplicated as of 2026-09-30, and the headline A/B comparisons above come from two of the seven failure types the paper tests across 168 scenarios. The title claim, that developer-written errors hurt the most capable agents most, is quantified in the per-model results: on expired credentials, the loss from the terminal command grew with model capability, from 18 points for GPT-5.5 to 35 for GPT-5.6 Sol, 39 for GPT-6 Sol and 69 for GPT-6 Astra; Astra’s loss exceeded GPT-5.5’s by 51 points (interval 29 to 72). A third comparison, on a missing permission, found naming the login tool raised recovery to 53%, against 28% for “Please re-authorize to continue.” The checklist below therefore stands on the measured deltas, on the independent tool-error studies further down, and on MCP spec error semantics.
The agent-consumable error: stable code, offending parameter, one named next step
An agent reading your error has exactly three questions: what class of failure is this, which input caused it, and what single action should I take. The payload that answers them:
| Element | What to ship | Why |
|---|---|---|
| Stable machine-readable code | A namespaced, versioned constant (auth/expired_token), not prose | Lets the agent branch deterministically; the logging guidance below depends on it |
| Offending parameter | The exact argument name and rejected value class | Converts “fix arguments” from guesswork into a targeted edit |
| One named next action | The literal tool or call to invoke, matched to the diagnosed failure | Naming the call moved recovery 6%→88% and 45%→84% in the measured scenarios |
The contrast case is the developer-oriented defaults measured in the two experiments: “Wait before retrying.” (no named action, 6% recovery), a terminal command (wrong actor, 45% recovery), and undifferentiated stack-trace prose (forces the model to extract the actionable fact from noise).
Note what the payload does not contain: a narrative of what went wrong internally. The tradeoff is real. Every internal detail you expose to help the agent recover is internal state you expose to every client, including ones you do not control. The three-element payload above is a reasonable ceiling: it reveals your tool names (already disclosed to clients by design) and one parameter name, nothing about your stack.
Where should this structure live? ToolRegistry frames every LLM tool call as structurally an RPC, a function name, JSON arguments, and a serialized result, and centralizes dispatch, schema generation, concurrency, and error recovery in a registry runtime. If your server is a bare API wrapper, the error translation layer is the piece you are currently missing: something must convert the upstream HTTP failure into the code-plus-parameter-plus-action payload, and that something belongs in one place in your server, not scattered across handlers.
Diagnose before you instruct: when recovery guidance backfires
The strongest temptation after reading the 6%-to-88% result is “add more guidance everywhere.” DARC (arXiv:2608.11772) measured why that fails. It identifies an “Intervention Mismatch”: a recovery signal chosen without diagnosis often cannot address the failure at hand. On the full admissible action set, DARC’s recovery prompt increased invalid actions per episode from 1.709 to 2.507; the prompt told the agent what the task required, but many resulting commands were not executable in the current state, so the extra guidance was spent on rejected actions. Under a guard-ranked view, where the same guidance was filtered to what was actually executable, invalid actions fell to 0.575.
The design rule: the named next action in your error must be executable from the state the agent is actually in. An error saying “retry search_flights” when the failure was an expired token sends the agent into a loop that burns budget. Diagnose first, then instruct, and if you cannot diagnose, say less. The credential experiment’s 82% recovery for a bare one-sentence prompt suggests that an honest “this failed” beats a confident wrong instruction.
One transfer caveat: DARC’s numbers come from embodied-action tasks, not MCP tool calls. Its direction (undiagnosed guidance backfires) is consistent with the credential experiment’s terminal-command result, but treat the specific magnitudes as evidence about the mechanism, not about your server.
A failure-injection protocol for your own server
Because the published headline deltas cover only two of the seven failure types tested, the actionable deliverable is a way to measure your own errors. Run this before each release that touches error paths:
- Enumerate failure classes. For each tool, list the error codes your payload can emit (rate limit, expired credentials, invalid parameter, upstream down, and so on).
- Inject each failure deterministically. Intercept at the wrapper layer and force each class, so the only variable is your error message.
- Run a fixed task set against each variant. Measure recovery rate: the fraction of tasks completed after the injected error. Also measure recovery cost in extra tool calls. MCP-Atlas fixes a 100-call budget per task and notes some failures stem from that ceiling rather than model capability, so every wasted retry consumes a real resource.
- A/B the wording. At minimum, test your current message against the three-element payload variant. The published experiments are your template: only the step’s content should differ.
- Test per model generation. The instruction-weighting report, if accurate, means wording tuned on one model can over-constrain the next. Re-run when you upgrade the client model.
- Watch for backfire. Count invalid or rejected actions after each injected error, DARC-style. If a guidance variant raises that count, the guidance is mismatched to the failure state; cut it back.
A variant you cannot recover from in testing is also a finding: some failures are terminal, and the right payload says so explicitly rather than inviting a retry loop.
Telemetry: per-call error codes as the debugging surface
Failure injection tells you how your errors perform before release; telemetry tells you how they perform after. The MCP Server Architecture Patterns paper recommends logging tool name, input hash, latency, output size, and error code per tool call, calling these logs the primary debugging surface for LLM misbehavior.
Two implications for the checklist. First, this is the concrete reason the error code must be stable and machine-readable: if your codes drift between releases, your per-call logs cannot be aggregated, and you cannot see that auth/expired_token recoveries dropped after a wording change. Second, log the code, not the prose. The message string is a rendering for the model; the code is the join key between your telemetry, your failure-injection results, and any client-side retry policy. The input hash matters because it lets you correlate repeat calls without logging raw arguments.
What better errors will not fix
Error wording carries real weight, but two findings bound how much.
Most random-draw servers fail before any error exists. From a 24,135-server MCP registry census, a random draw of 400 npm/stdio servers saw only 48.8% complete the initialize handshake, against 66.7% for a hand-curated frame measured with the same instrument. The dominant failure in that census was servers that never start at all (37.5%), ahead of missing credentials (13.3%). A server that never starts emits no error message. The same census found hard JSON Schema conformance essentially total (zero fatal violations across 2,766 advertised tools among the 195 servers that ran) while optional safety annotations were omitted on 58.8% of tools. The ecosystem’s structural problems are liveness and annotation, not schema syntax; error wording sits downstream of all of them. If you are triaging a server portfolio, a keep, fix, or drop rubric for tool servers addresses the earlier decisions.
Many agent failures are not error-driven. MCP-Atlas, a benchmark against real MCP servers, found that the dominant failure for current agents often happens after tool access: understanding the task, deciding whether enough evidence has been gathered, and synthesizing the final answer. For o3 Pro, 57.6% of diagnosed failures were classified as tool-related, but 40.1% of its failed trajectories contained no tool invocation at all despite tasks requiring external evidence. Improving your error messages fixes the slice of failures where an agent hit your error and mishandled the recovery. It does nothing for the agent that never called you.
Verdict
Ship MCP tool errors as a stable machine-readable code, the offending parameter, and one explicitly named next action matched to the diagnosed failure state. Validate the wording with failure injection before release, measure recovery rate and invalid-action counts per variant, and log per-call error codes so regressions are visible. Where you cannot diagnose the failure, emit less guidance, not more: DARC measured undirected recovery prompts raising invalid actions from 1.709 to 2.507, and the credential experiment showed developer-oriented instructions (45% recovery) losing to near-silence (82%).
Hold three limits in view. All recovery-rate deltas come from one researcher-run paper, unreplicated, whose headline A/B comparisons cover two of its seven failure types. The 88% headline number from that experiment describes rate-limit recovery specifically, not aggregate task success, and MCP-Atlas shows the dominant agent failures occur after tool access regardless of error quality. And for the 37.5% of randomly drawn registry servers in the census that never start, there is no error message to redesign. Fix liveness first, then fix what your errors say.

Join the discussion
Share a useful perspective or ask a question about this article.