The Context Window

MCP Context Window Budget Management

MCP sessions waste tokens on tool definitions before users even ask questions.

Contributing Editor · · 12 min read
Cover illustration for “MCP Context Window Budget Management”
Model Context Protocol · September 4, 2026 · 12 min read · 2,589 words

MCP's real engineering problem shows up after the tools are already linked to the model. The harder question is deciding what fraction of a fixed token budget those tools get to eat before a human even types a question. Model Context Protocol is an open standard, model-agnostic, meant to give AI systems one uniform way to talk to external tools and data sources instead of a different custom wire format for every vendor. Unlike REST, which is stateless and organized around resources, MCP is stateful and session-based, built on JSON-RPC 2.0. That distinction sounds like plumbing trivia. It carries real consequences.

A REST API doesn't remember you between calls. An MCP session does, and that memory has to live somewhere: the context window. The host application, whatever sits between the model and the tools, has to manage that window: deciding when to call a tool, routing the request, and feeding the result back into the conversation. Two terms get flattened together constantly in discussions of this problem, and the flattening causes real confusion. The context window is the token budget for a single inference call; it sets capacity, full stop. The context layer is everything upstream of that, the infrastructure deciding what actually earns a seat in the window. Conflate the two and you start treating a governance problem like a storage problem, which is a bit like blaming a full parking garage on the pavement instead of the valet who keeps waving cars in.

This isn't a niche concern anymore. MCP's TypeScript and Python SDKs have both crossed the billion total downloads threshold, and by mid-2026 the Tier 1 SDKs were pulling close to half a billion downloads a month. In December 2025, Anthropic handed MCP over to the Agentic AI Foundation under the Linux Foundation, co-founded with Block and OpenAI, a fairly clear signal that governing the protocol is now everyone's problem, not one company's roadmap item. The rest of this piece walks through why that governance question is where MCP implementations succeed or quietly fall apart.

How tool definitions colonize the context window before a session begins

Here's the default behavior nobody warns you about: when an MCP session starts, every tool definition from every connected server loads into context immediately. No filtering, no waiting for the model to ask. It all just shows up, uninvited.

The cost per tool isn't trivial either. A properly documented MCP tool, with parameter descriptions written the way they're supposed to be, typically runs from a few hundred tokens to well over a thousand. Connect a couple dozen tools and that arithmetic turns hostile fast. One developer audit from 2025 found that tool definitions alone consumed tens of thousands of tokens before a single user query got typed, across just a handful of connected MCP servers. A third of the budget, gone, before the conversation even starts. In more extreme setups, benchmarks have shown that traditional MCP servers with very large toolsets can burn more tokens on schema definitions alone than Claude Code's entire maximum context window allows. That's a server configuration a team might actually ship.

And the schema preload is only one leg of the tax. Every tool call returns output, and that output accumulates across the session, an ongoing toll rather than a one-time fee. Layer on top of that the conversation history itself, growing turn by turn and competing for the same fixed pool of tokens. Three separate cost centers, all drawing from one account.

There's a second-order effect worth naming directly: bloat doesn't just cost money, it costs accuracy. Models get worse at picking the right tool when they're wading through pages of irrelevant tool documentation, and this happens before the "lost in the middle" dynamic (more on that shortly) even enters the picture. Ask a practitioner running MCP servers at scale for a rule of thumb, and a common answer is somewhere around ten to fifteen active tools at a time. That's a heuristic, not a line in the spec, but it tells you something about where experienced teams have found the ceiling actually sits.

Why expanding the context window does not fix the budget problem

The obvious fix, when your tool definitions are crowding out the actual conversation, is to just get a bigger window. More tokens, more room, problem solved. The research on this point tells a more complicated story.

Liu et al. (2024) documented what's now called the "lost in the middle" effect: model accuracy on tasks like multi-document question answering and key-value retrieval follows a U-shape across the context window. Information at the very start or the very end gets found reliably. Information buried in the middle gets missed, sometimes badly. This wasn't a quirk of one model either; the pattern held across six different model families tested, including GPT-3.5-Turbo, GPT-4, Claude 1.3, LongChat-13B, MPT-30B, and Cohere Command. The RULER benchmark pushed this further, testing 17 long-context models, and found recall degraded as input length grew in every single one. A bigger window doesn't remove the middle; it just moves the middle further away and gives the model more room to lose things in it.

Anthropic's own context engineering research, published in September 2025, landed on the same counterintuitive finding: accuracy goes down as the context window gets bigger, not up. The mechanical answer sits in how self-attention works: its complexity scales quadratically with input length, so a longer context isn't just "more to read," it's a direct computational penalty on both latency and accuracy at the same time. There's one notable exception worth flagging: Google's Gemini 2.5 Flash holds up well on needle-in-a-haystack retrieval regardless of where the needle sits. But that result is specific to simple factoid lookup, not the messier, multi-tool, multi-turn scenarios MCP agents actually operate in. The broader lesson holds anyway: the fix was never going to be dimensional. If the problem is what gets into the window and in what order, no amount of extra square footage solves it. The fix has to be architectural.

Diagram: Bigger Windows, Worse Recall: The U-Shape Problem. Visualizes: Visualize the 'lost in the middle' effect on model accuracy across the context window.

The practical approaches teams use to govern tool context before the spec catches up

If bigger windows don't help, what does? Mostly, keeping things out of the window until they're actually needed.

Lazy loading, sometimes called on-demand discovery, keeps tool schemas out of context entirely until the model signals it actually wants to use one. Claude Code implements this in production with a search tool that surfaces relevant tool definitions only when asked for them, turning tool discovery into a routing decision the host controls rather than a tax charged automatically at session start. Some third-party "Lazy Load" MCP servers claim very large reductions in context token consumption from this approach; the exact figures are vendor-reported, so treat them as directionally credible rather than gospel, but the mechanism itself checks out.

Progressive disclosure works on a similar principle, staged rather than binary: reveal that a tool exists first, and only fetch its full parameter schema once the model looks ready to actually call it. Production teams using this pattern have reported large drops in token usage without losing accuracy in tool selection, which suggests the underlying idea, defer detail until it's needed, holds up outside the lab.

One team took a more structural approach and collapsed an entire data layer into just two MCP tools: a schema introspection tool and a query execution tool, using GraphQL as a compression layer. Instead of dozens of narrow, single-purpose tools each carrying their own documentation weight, two general-purpose tools absorbed the same functionality for a fraction of the token cost.

Prompt caching and batching help too, but only if they're designed into the pipeline from the start rather than bolted on after the fact once someone notices the bill. And for long-running sessions where the conversation history itself becomes the threat, research into prompt compression techniques (DPEPO-style approaches, for instance) has shown that prompts with very large token counts can shrink to a fraction of their original size while the agent's actual performance holds steady. What ties all of these together is where the decision sits: it's the host making the call about what enters the window, when, and in what shape. The model itself has no idea what got left out. It just sees what it sees.

What research-grade tool retrieval systems show is actually achievable

Diagram: Tool Retrieval at Scale: Token Cost vs. Accuracy. Visualizes: Show a comparison of three research-grade tool retrieval approaches against a static baseline, along two dimensions: token reduction and tool selection accuracy.

The academic work on this problem is further along than most production systems, and the gap is worth sitting with for a second.

RAG-MCP, published in May 2025, applies retrieval-augmented generation to the tool selection problem itself: a semantic index picks out the most relevant tool descriptions for a given query, and only those descriptions get inserted into the prompt. The model never sees the full catalog. Results: prompt tokens dropped by more than half, and tool selection accuracy rose from a 13.62% baseline to 43.13% on benchmark tasks, more than tripling accuracy while shrinking the context footprint at the same time. Worth noting the scale this is solving for: as of April 2025, the MCP server ecosystem had grown to more than 4,400 listed servers. Retrieval at that scale isn't a nice-to-have, it's table stakes.

Speakeasy's Dynamic Toolsets, benchmarked in November 2025, take a related but distinct approach: compress a large toolset down to three meta-tools, and let the model search for the specific tool it needs using natural language and embeddings-based semantic search, then describe and execute from there. The numbers reported are striking: up to 160 times token reduction against static toolsets, a 96% cut in input tokens, and roughly 90% reduction in total token consumption on average, all while keeping a 100% task success rate in testing. But there's a real cost on the other side of that ledger. Dynamic toolsets need two to three times more tool calls than static ones, and Speakeasy measured about a 50% increase in execution time. Efficiency in tokens, paid for in speed.

MCP-Zero, presented at NeurIPS 2025, goes further still with a three-part system: an Active Tool Request mechanism where the model generates its own structured tool requirements, Hierarchical Semantic Routing that matches queries to servers first and tools second, and Iterative Capability Extension for building cross-domain toolchains. Tested against 308 servers and 2,797 tools pulled from the MCP official repository, roughly 248,000 tokens of schema in total, it achieved accurate tool selection while cutting token consumption by 98% on the APIBank benchmark, without giving up accuracy. Maybe the most useful finding, given how agents actually get used, was multi-turn resilience: accuracy dropped by no more than 3% going from single-turn to multi-turn scenarios, where standard methods degraded sharply as conversation history piled up. Most production agents are multi-turn by nature, so a benchmark that only tests single-turn performance is flattering approaches that would fall apart the moment a real conversation started.

Three different systems, three different mechanisms, one shared idea: externalize the tool selection decision. Let the host retrieve, filter, or route before the model ever sees the full menu, instead of asking the model to sort through everything at once.

How the MCP specification itself has started encoding these lessons

The spec has been catching up to all of this, and the changes read less like feature additions and more like lessons learned the hard way, now written into the rulebook.

The June 2025 spec update (v2025-06-18) introduced a resource_link type: tools can now point to a URI instead of inlining full content directly into the response, and the client fetches or subscribes to that URI only when it's actually needed. That's lazy loading built into the protocol itself, arriving after third-party servers had already worked out the same pattern on their own. The July 2026 update went further with cacheable list results: list responses now carry cache hints and deterministic ordering, so clients can cache tool catalogs and keep upstream prompt caches stable across reconnects. That directly targets the schema-preload cost described earlier, specifically at the moment of reconnection, which is exactly when that cost tends to resurface.

Claude Code implemented a search tool as its production approach to on-demand tool loading, and reported measurements showed a meaningful drop in context consumption as a result. It stands as the clearest vendor-reported data point for how much this kind of lazy loading actually saves in a live product, distinct from a benchmark paper.

The most structurally significant change, though, is the one buried in the July 2026 spec: MCP introducing structural changes aimed at reducing persistent session-state accumulation, one of the root causes behind the compounding bloat that long-running agents run into over time. It was also a change that developers building on the protocol had long pointed to, which says something: the people building this stuff had already identified session state as a budget threat well before the spec formally addressed it. Line up all four of these changes, resource_link, cacheable lists, tool_search, and the stateless core, and a pattern emerges. The spec is deliberately moving context governance decisions out of the model's limited attention and into the protocol and client layers where they can actually be managed.

What a well-governed MCP implementation actually looks like in production

So what does this look like once it's actually running? The single decision that shapes everything downstream is how the host allocates, tracks, and prioritizes what enters the context window, and when. Connectivity is the easy part; governance is the job.

A practical way to think about this breaks into four layers. First, catalog management: keep an external index of tool descriptions and never load the full catalog into context wholesale. Second, intent routing: retrieve the schemas relevant to the current query before the model sees anything else, which is exactly what RAG-MCP and MCP-Zero demonstrate is achievable at real scale, not just in a lab notebook. Third, progressive execution: only load full schemas when a tool is actually about to run, leaning on resource_link and cacheable list results wherever the spec supports them. Fourth, session hygiene: use stateless protocol modes when they're available, and compress or summarize old conversation history before it starts crowding out what's actually relevant right now.

None of this is free, and pretending otherwise would be dishonest. Dynamic tool retrieval adds latency, more tool calls, longer execution windows, the tradeoffs Speakeasy's benchmarks quantified directly. Whether that tradeoff is worth making depends on what's actually the bottleneck for a given system: token cost, selection accuracy, or raw response speed. Pick the wrong priority and optimize for it anyway, and the fix will feel worse than the disease.

It's also worth pointing out that tool descriptions carry real weight as the actual semantic surface the retrieval layer searches against, rather than functioning as documentation afterthoughts written once and forgotten. A poorly worded tool description is invisible to intent routing no matter how sophisticated the host's retrieval system is. That makes context budget governance an ownership question as much as a technical one, something that needs a deliberate owner at the architecture or product leadership level, because nobody else is positioned to see the whole budget at once.

And this isn't unique to MCP tool servers, either. Anywhere AI agents get assembled out of multiple tool-connected workflows, the same token math applies, whether that's customer support routing or AI-assisted content pipelines juggling research tools, style guides, and publishing systems all at once. The failure mode looks identical everywhere it shows up: bloat, confusion, and a cost curve that bends the wrong direction. Governance architecture is what separates the systems that scale from the ones that quietly buckle under their own context.

Sources

  1. blog.modelcontextprotocol.io
  2. getunblocked.com
  3. sparkco.ai

More in Model Context Protocol