MCP Tool Call Error Handling Patterns

MCP tool call errors are not random noise; they fall into predictable categories, and each category demands its own recovery pattern. Get the categorization right and an AI agent recovers on its own, but get it wrong and every failure looks like a five-alarm fire, whether it's a typo in a parameter or a database that fell over.
Model Context Protocol, the open standard Anthropic introduced in late 2024, gives large language models a common way to talk to external tools and data. Three roles make up the system: the host is the user-facing application, the client is the protocol layer living inside that host, and the server is whatever exposes the tools, resources, and prompts. Each client keeps a one-to-one session with a specific server; it formats requests, tracks session state, and processes whatever comes back. Under the hood, MCP extends JSON-RPC 2.0, the lightweight standard for calling procedures over JSON, and adds LLM-specific concerns on top: schema validation, context continuity, handling the fact that models are sometimes wrong about what they're asking for.
A single tool call crosses several boundaries on its way from thought to result. The host builds the request, the client serializes it, the message travels across some transport (HTTP, stdio, whatever the deployment uses), the server validates it, the actual tool runs, and the output makes the same trip in reverse. Each boundary is its own failure surface. That's really the whole premise of this piece: errors aren't scattered at random across the system, they cluster at specific points, and once you know where, you know what to do about them.
Everything below reflects the current stable spec, dated 2025-11-25, along with the 2026-07-28 release candidate, which changes some error code behavior worth knowing before it locks in.
The two-path error model that makes MCP different from ordinary API failures
MCP splits failure into two lanes, and mixing them up is probably the single most common mistake in early implementations.
Protocol errors mean something broke in the message exchange itself: malformed JSON, a method that doesn't exist, parameters that don't match the schema. These come back as a JSON-RPC error object. Tool execution errors mean the message got through fine, the tool ran, and the tool's own logic failed. These come back as a normal JSON-RPC result, just with a field called isError set to true.
The reasoning behind the split is worth sitting with for a second. Execution errors go back to the model as plain, readable content that it can reason about and try to fix. Protocol errors mean the request itself was broken, and there's not much a model can do about a request it didn't correctly understand how to build in the first place; that's a client bug, not a reasoning gap.
Here's where it goes wrong in practice. Wrap an API timeout in a -32603 Internal Error, and the model never sees a message it can act on; it just sees "something broke," full stop. Flip it the other way and return a tool execution error for invalid parameters that were the client's fault to begin with, and now the model is trying to "fix" its own reasoning for a mistake that happened somewhere else in the stack entirely. Neither is fatal on its own, but do it often enough across a codebase, and the agent starts treating every failure as unrecoverable, which defeats most of the point of building an agent.
Protocol-level errors and what each code actually signals
The standard JSON-RPC codes carry over into MCP mostly unchanged. -32700 means invalid JSON came in, -32600 means the JSON was valid but wasn't a proper request, -32601 means the method or tool doesn't exist, -32602 means the parameters were wrong for the tool being called, and -32603 is the internal error catch-all.
Most implementations never write these by hand; the MCP SDK handles them and surfaces them automatically. What's changing is the range these codes live in, and the 2026-07-28 revision moves some furniture around that's worth flagging now.
The -32000 through -32019 range is now legacy. New implementations shouldn't allocate codes there, and receivers shouldn't assume any particular meaning for codes in that band, with one exception carved out for -32002. The -32020 through -32099 range is now reserved specifically for MCP-defined codes: -32020 for HeaderMismatch, -32021 for MissingRequiredClientCapability, -32022 for UnsupportedProtocolVersion. And -32002, which used to mean "resource not found" under the 2025-11-25 spec and earlier, gets replaced by -32602 going forward. Clients still need to accept -32002 from servers running the older spec, so this is additive complexity for a while, not a clean swap.
The practical rule holds regardless of which code shows up: when a protocol error lands, fix the client's request construction or the transport setup. Don't hand it to the model and hope for the best; the model didn't build the broken request, the client did.
Worth noting that this whole area is still under active debate in the MCP community, and treating today's code ranges as permanent would be a mistake.
How to structure tool execution errors so the model can act on them
The wire format for a tool execution error looks, from the transport's perspective, like a success. The JSON-RPC response comes back fine; inside it sits a result object with isError set to true and a content array holding the details. Transport sees success, while application sees failure, and that gap is the whole point.
What goes inside that content array is where the actual design work happens, and there are roughly three ways to get it wrong before getting it right. A raw stack trace is noise; an LLM has no reliable way to extract "what should I do differently" from a Python traceback. A generic string like "Operation failed" is arguably worse, since it tells the model nothing about whether to retry, change its inputs, or give up. What works is a precise, human-readable description of exactly what went wrong and why.
Take a Pydantic validation error as the concrete case. Parse the list of failed fields and produce something like: "Validation Error: Argument 'quantity' expected an integer, but received 'five'." Now the model can see the type mismatch directly and re-invoke the call with 5 instead. It's a small rewrite of the error message, and it turns a dead end into a next step.
This connects to a deliberate design choice in the MCP spec: input validation errors should come back as tool execution errors, not as -32602 protocol errors, specifically so the model has a chance to correct them. That's a deliberate design choice, not an accident of implementation. The isError flag is the envelope; the message inside it is the actual error handling work, and it's worth remembering that the envelope alone does nothing for the model. Structured metadata about the error category and whether a retry makes sense belongs in that same response object. The next section explains why that categorization carries so much weight.
Four error categories and why each needs a different recovery response
Every tool failure sorts into one of four buckets, and the model's next move should follow from the bucket, not from guesswork. Without category metadata attached to the error, the agent is essentially trying things until something works. With it, recovery becomes close to deterministic.
Transient errors cover timeouts, service unavailability, rate limits: cases where the underlying system is briefly unreachable but the request itself was fine. Recovery here is a delayed retry, and the error response should signal that retrying is appropriate. Validation errors are parameter format or type mismatches, the case the Pydantic example above covers directly. The model corrects its inputs and calls again — but only with corrected inputs, not the same ones.
Business or logic errors are a different animal entirely. The tool ran cleanly, no crash, no timeout, but the result is a domain-level failure: "account not found," "order already closed." Retrying with identical inputs will fail every single time, so the error response should signal that retrying is futile and the model needs to replan rather than retry. Permission errors, the fourth bucket, cover access denied and authentication failures. There's no parameter fix or retry that solves this from the model's side, and the right move is escalation — either to a human or to a coordinating process — because the agent genuinely cannot resolve this on its own.
For multi-agent setups specifically, the architectural principle worth following is to keep recoverable failures local and only surface genuinely non-recoverable failures upward, ideally bundled with partial results and a log of what was already attempted. Skip this filtering and higher-level processes drown in noise from failures that should have been handled further down the stack.
Retry logic, backoff, and circuit breakers for transient failures
Here's a scenario worth walking through. A tool call fails, the host detects the failure, and it retries, then retries again, and again, in rapid succession, because nothing in the system told it to slow down. Without server-side retry intelligence, a database that's already struggling under load gets hit with a wave of near-simultaneous repeat requests, which is a strange way to fix an outage by making it worse.
Exponential backoff with jitter is the standard fix. Each retry waits longer than the last, typically doubling the delay each time, and jitter adds a small randomized offset to that delay so that many clients failing at the same moment don't all retry in perfect synchronization and slam the server together. Fixed-delay retries, where every attempt waits the same interval, should be treated as an anti-pattern anywhere more than one client might be hitting the same server.
The same logic applies to polling for long-running tasks, which some MCP deployment patterns treat as a first-class concern. Poll backoff should scale exponentially too, not sit at a fixed interval, and capping it at some reasonable maximum keeps traffic from spiking during a server's busiest windows.
Circuit breakers sit one layer above retries. Where a retry says "try again after waiting," a circuit breaker says "stop trying altogether for now." Once a tool's failure rate crosses some threshold, the circuit opens, and further calls fail immediately instead of waiting out a timeout that was probably going to fail anyway. Layered together, the full resilience stack usually looks like rate limiting, then bulkheads, then circuit breakers, then retries, with each layer catching a different kind of failure mode.
This gap shows up across MCP client libraries generally — the absence of built-in retry and circuit breaker logic means transient failures can surface as hard errors, which is a fairly clean signal that the problem isn't specific to any one project. The division of labor that falls out of this: retry logic belongs in the client, circuit breaker logic belongs at the session or agent level, and neither is something the server should be expected to enforce on the client's behalf.
Using structured error feedback to close the model's self-correction loop
An unhandled exception that takes down the server ends the conversation, full stop. A structured error message the model can actually read, on the other hand, becomes one more piece of input for the next reasoning step. That difference is really the whole argument for building error handling carefully instead of bolting it on at the end.
The loop in practice runs like this: a tool fails, a structured error comes back with isError true and a precise message, that message lands in the model's context, and the model generates a corrected call. It works reliably for validation errors and a fair number of logic errors, though the whole thing depends entirely on whether the error message actually says something useful, since garbage in means garbage reasoning out.
In practice, models pick between a range of recovery strategies depending on what failed: retry, replan, swap in a different tool, or hand off to a human. What makes that selection non-arbitrary is the error category and the message content; strip those away and the model is choosing blind.
One failure mode worth naming directly: even with temperature set to zero, models still generate incorrect tool-call arguments, whether that means wrong file paths, mismatched parameter types, or coordinates that don't correspond to anything real. No amount of prompting eliminates this completely, because the model can't guarantee every invocation is correct before it happens. Structured MCP error feedback, whether it's a file-not-found message, a type error, or a flag for an invalid region, is what lets the agent adjust its next move instead of just stalling out.
An architecture worth watching separates self-healing from LLM-driven recovery entirely. A Tool Health Monitor layer can reroute around a known-failing tool algorithmically, no model call required, just finding the next viable path through the available tools. The LLM only gets consulted when no algorithmic alternative exists, which preserves inference budget for the failures that are genuinely ambiguous rather than burning a model call on something a lookup table could have solved. Compare that to the more familiar ReAct-style loop, where every tool failure triggers several rounds of LLM reasoning just to identify an alternative that a simpler system might have found instantly.
The upshot: the text inside an error message isn't documentation written for the next developer who opens the file. It's runtime instruction for a reasoning system that's going to act on it in the next few hundred milliseconds, and it should be written that way.
Long-running tool calls and how the Tasks extension changes error handling
Synchronous tool calls block until the tool finishes, which is fine when the tool takes 200 milliseconds and a real problem when it takes 90 seconds. The Tasks extension, introduced as experimental in the 2025-11-25 spec, exists to fix exactly that.
Tasks flip the model from "call and wait" to "call now, fetch later." The client submits a task and gets a task ID back right away, then polls tasks/get to check status. The tool itself runs asynchronously on the server, on its own schedule, with no connection held open the whole time.
This opens up failure surfaces that plain synchronous calls never had. The task submission itself can fail, which is a protocol error at the moment of submission, same as any other. The task can also fail during execution, except now that failure only surfaces the next time the client polls, not at the moment the call was made. And polling itself can fail, independent of whatever the task underneath is actually doing, since a network hiccup or an overloaded server can break the poll without the task having failed at all.
The design response follows the same principles laid out earlier in the piece, just applied to a new shape. Poll failures get retried with exponential backoff, independent of whatever the task's status turns out to be. Task execution failures still use the isError pattern inside the task result, same two-path model as any synchronous call. And tasks need a hard maximum lifetime along with a retry cap on poll failures, because an agent stuck polling forever for a task that will never finish is just a slower, quieter version of the retry storm from a few sections back, and just as capable of leaving a system stuck.


