MCP Sampling Requests and Model Delegation
Servers can borrow the client's model, but the spec is already reworking how.

MCP sampling lets a server borrow the client's language model instead of buying its own. That's the whole idea, and as of the 2026-07-28 spec, it's also a deprecated one, with the specification moving toward approaches that keep the idea and rework the underlying wiring. This piece walks through how sampling works, where it breaks down in practice, and why its replacement isn't a reversal so much as a plumbing fix.
Start with the standard setup, because sampling only makes sense once you see the wall it was built to get around. A normal MCP exchange is polite and a little dumb: client sends a request, server runs a tool or hands back some context, client gets the result and moves on. The server in this picture is a vending machine. It has tools, it has data, but it has no opinions about any of it. Fine, until the server is doing something like code review, or security analysis, or a multi-step research task, where it needs to stop halfway through and actually reason about what it just found. Under the plain request-response flow, there's no clean way to pause and think. The server either needs its own model access (an API key, a bill, a whole separate integration) or it stays deliberately stupid and shoves every decision back to the client to sort out. Sampling exists to ask a more useful question: what if the server could just borrow the client's model instead of buying its own?
What sampling is: the server asks the client to run the model on its behalf
Sampling is the protocol-level mechanism that reverses the usual direction of an LLM call. Instead of the server making its own request to OpenAI or Anthropic or whoever, it asks the client to do it. The client already has the model access, the API key, the billing relationship, so the server just piggybacks on that. Less "the server got smart," more "the server found a hall pass."
The practical upside is straightforward: a server can use a language model without ever holding a credential for one. Cost and access stay entirely on the client's side of the fence. A server author can write reasoning into their tool logic (summarize this, flag this, classify this) without turning their MCP server into a product that needs its own OpenAI billing account and rate-limit headaches.
Sampling has a cousin worth separating out: elicitation. Both involve the server reaching back mid-call for something it doesn't have. Sampling asks the model. Elicitation asks the human user directly, for something like a missing parameter or a confirmation. Same shape, different target. One borrows reasoning, the other borrows judgment.
None of this is magic. The server delegates a specific, bounded task; it doesn't seize control of the client. The client decides what actually happens, and that distinction matters a few sections down, when the spec starts leaning hard on human approval as a required checkpoint.
The sampling/createMessage request: what the server sends and what the client controls
The actual method is sampling/createMessage, a JSON-RPC request the server fires off to the client. It carries the messages to send the model, an optional systemPrompt (something like "You are a security analyst"), a maxTokens ceiling, and the usual generation knobs: temperature, stopSequences.
Here's the part that trips people up on first read: none of those parameters are commands. They're requests, full stop, and the spec is explicit that the client can override any of them. Client shortens maxTokens because someone's on a budget this month? Fine. Client rewrites the systemPrompt because the original one was going to make the model act weird? Also fine. The server proposes, the client disposes, and that asymmetry isn't an oversight, it's the entire safety model.
What comes back is more informative than a plain completion would be. The response includes the role, the content, a stopReason, and, critically, which model actually ran the request. That last part matters more than it sounds like it should, because the server asked for a model but doesn't get to assume it got the one it named.
Requests aren't limited to text either. Content types include text, base64-encoded images, and base64-encoded audio, so a server can ask for a multimodal completion if the client supports it. None of this works unless the client declared sampling support during the initialization handshake. No declaration, no sampling. The handshake is the permission slip.
How the client picks a model when the server can't name one directly
A server writing a sampling request has a problem that sounds small but isn't: it can't just say "use gpt-4o" or "use claude-3-opus," because the client might not have access to either. Naming a specific model bakes in an assumption about the client's provider that the server has no business making.
The spec's answer is modelPreferences, an object that separates two kinds of hints. First, three numeric priorities on a 0 to 1 scale: costPriority, speedPriority, and intelligencePriority. These are abstract dials, not model names, the equivalent of telling a barista "something strong" instead of naming a bean. Second, a list of advisory hints, name substrings like "claude-3-sonnet" with fallbacks like "claude," evaluated in order but not binding.
The spec itself gives an example worth sitting with: a client that doesn't have Claude access at all might see the "sonnet" hint and map it to something like gemini-1.5-pro based on roughly comparable capability. The server's intent (mid-tier model, balance cost and quality) survives the translation even though the literal model name doesn't. The server describes what it wants in terms of tradeoffs, and the client figures out which actual model on its actual provider account satisfies that.
Final say always sits with the client. Hints are advisory, priorities are advisory, and none of it is enforceable from the server's side. That's by design: it keeps server code from hardcoding a dependency on a provider it may never touch directly.
The human-in-the-loop requirement built into the protocol
The spec doesn't hedge much on this one: there SHOULD always be a human able to deny a sampling request. Not "could." Should. And it names two specific checkpoints where that denial has to be possible.
Before the prompt goes to the model, the client is supposed to show the user what's about to be sent, giving them a chance to edit it, approve it, or kill it outright. After the model responds, the client is supposed to show the user the completion before it goes back to the server, again with room to edit, flag, or ask for a redo. The spec also asks that client UIs make this review step easy and intuitive to use, not something users are likely to skip past.
Why build a review gate into a protocol meant to enable autonomous agent behavior? Because the two goals sit in real tension with each other, and the spec doesn't pretend otherwise. A fully hands-off agentic loop and a mandatory human checkpoint pull in opposite directions, and the "SHOULD" instead of "MUST" phrasing is the spec's quiet acknowledgment of that friction. The intent is clear enough anyway: the server should get back an approved response, not a raw, unreviewed model output smuggled through the pipe.
Where security practice lags behind the spec's intent
Good rule on paper. The less flattering part is what actually happens in shipped software, and here the record is not close to good.
Security research has documented a specific gap, sometimes labeled "Sampling Without Origin Authentication," and it's dumb in the way a lot of real vulnerabilities are dumb: the client doesn't distinguish, on screen, between a prompt injected by the server and one typed by the human user. Three major MCP host implementations examined in that research, Claude Desktop, Cursor, and Continue, all showed the same gap. None flagged sampling-derived messages as coming from a server rather than a person. A user approving a sampling request can't easily tell whether they're reviewing their own words or a server's, which makes the human checkpoint from the previous section closer to theater than to a control.
Zoom out and the picture doesn't improve. Independent scans of the public MCP server ecosystem have found exploitable flaws in a wide range, from 30% to 82% of public servers depending on the scan, and only 8.5% using OAuth for authentication. Adoption kept climbing well ahead of any of that getting fixed: over 10,000 active public servers and more than 97 million monthly SDK downloads. Security maturity did not scale at the same rate as popularity, which is the pattern that shows up almost every time a protocol gets popular faster than anyone can audit it.
The CVEs are real, not hypothetical. CVE-2025-6514 alone reached hundreds of thousands of downloaded environments. CVE-2025-49596, tied to MCP Inspector, was pulling tens of thousands of weekly downloads. CVE-2025-54136 is in the same class of known MCP-related vulnerabilities, collectively reaching large numbers of developer environments.
There's also a quieter risk baked into the architecture itself: a server that calls sampling recursively without a hard depth limit can rack up unbounded model costs off a single malformed input. And the protocol has no built-in concept of semantic attenuation, meaning there's no native way to say "this sampling call only gets read-only scope." Access is binary. You either can sample or you can't, nothing in between.
None of this means the human-in-the-loop idea is bad design. It means the clients implementing it haven't caught up to what the spec actually asks for, and a checkpoint only works if the person at it has enough information to make a real decision. Right now, in several major clients, they don't. That gap is worth remembering later in this piece, because it's part of the argument for why sampling's replacement had to change more than just the plumbing.
What the November 2025 spec added: tool calling inside sampling requests
The 2025-11-25 revision closed a notable gap: no support for calling tools inside a sampling request. That's a bigger deal than it sounds, because tool calling is more or less the load-bearing wall of agentic behavior. Without it, sampling could think, but it couldn't act.
The update let sampling requests carry tools and toolChoice parameters, so the model running on the client side can call tools as part of generating its response, with toolChoice controlling how much freedom the server wants to hand over. The model can take the results of tool calls and keep going within that same exchange rather than needing a fresh round trip for every step.
Put together, that produces something close to a full agentic loop contained inside one delegated call: server hands off a task, client-side model reasons about it, calls whatever tools it needs, processes what comes back, and returns a finished answer, with the human approval gate from the previous section still sitting at each step. The same release also soft-deprecated includeContext in favor of more explicit capability declarations, tightening up context handling for exactly this kind of multi-step flow.
Worth noting this wasn't an isolated patch. The 2025-11-25 release was a substantial one in MCP's history. It was part of a broader push to make MCP behave like production infrastructure rather than a demo. Tool-calling inside sampling was one piece of a broader push to make MCP behave like production infrastructure rather than a demo.
Practical patterns where sampling earns its complexity cost
So where does anyone actually reach for this? A few patterns show up often enough to count as the standard playbook.
Code review is the obvious one: a server midway through reviewing a pull request realizes it needs more context than it started with, and instead of trying to reason about it internally, it asks the client's model to summarize the recent change history. The server stays small and dumb about the actual reasoning; the thinking happens on borrowed compute. Security analysis works the same way, with a server sending a sampling/createMessage request carrying a systemPrompt like "You are a security analyst," so the reasoning becomes part of the server's logic without the server ever needing to be an AI product in its own right. Research summarization follows the pattern too: a server with a tool that fetches raw pages from Wikipedia gathers the material, then hands it to sampling to turn into something coherent, splitting the job cleanly between fetching (server) and synthesis (client model).
On the implementation side, frameworks like FastMCP take care of some of the tedium that would otherwise make this fiddly. Schema generation, prompt engineering for structured output, response validation, that scaffolding gets handled so a server can return a validated Pydantic object from a sampling call without hand-rolling all of it.
The thread running through all of these: sampling fits where a server owns domain logic or data access but needs a reasoning step it has no business owning permanently. It's a loan, not a handoff. Each case quietly benefits from the two mechanisms covered earlier, the model-preference system picking a reasonable model for the job, and the human checkpoint (theoretically) catching a bad security finding before it gets treated as gospel.
The 2026-07-28 specification deprecates sampling and explains why the design had a structural limit
Here's the turn the rest of this piece has been building toward: as of the 2026-07-28 specification, Sampling gets formally deprecated, alongside Roots and Logging. And the reasoning behind it says more about distributed systems than it does about the idea of sampling itself.
Deprecated doesn't mean dead on arrival. The formal lifecycle policy requires at least a twelve-month window between deprecation and eligibility for removal, so nothing currently running on sampling breaks overnight. There's a narrower exception: an expedited removal exception is available in certain circumstances, though it must still provide at least ninety days between a feature becoming Deprecated and its earliest removal, but even then a minimum of ninety days has to separate deprecation from removal. New projects, though, are told plainly not to build on sampling going forward. Existing ones get a runway, not a cliff.
Why deprecate a mechanism that was, by most accounts, working as designed? The answer is architectural, not a change of heart about the idea. Sampling and elicitation both depended on a persistent connection between client and server: the server pushes a request down an active connection while the original tool call sits there, pending, waiting for the answer to come back. That's elegant when there's one server instance and one long-lived connection. It falls apart the moment anyone tries to run that server the way most backend services actually get run today: stateless, horizontally scaled, sitting behind a load balancer with no memory of which instance handled the last request. Supporting sampling at that scale means sticky routing, shared coordination state, long-lived infrastructure, or pinning a client to one specific instance forever. What made sampling clean in a single-box setup turns into a real deployment tax at scale, and that tax is the actual reason it's going away, not the security gaps covered above (those are a separate, if related, argument for moving fast).
The SDKs are already reflecting the shift. The Python, TypeScript, Go, and C# SDK beta releases mark sampling, along with roots and logging, as [Obsolete], pointing implementers toward the replacement pattern the spec now recommends outright: call the LLM provider's API directly instead of routing through the client.
How Multi Round-Trip Requests (MRTR) preserve the interaction model without the stateful stream
The replacement is called Multi Round-Trip Requests, MRTR, tracked as SEP-2322, and it swaps out server-initiated sampling, elicitation, and roots/list requests for something that doesn't need a connection to stay open between turns.
The mechanism: instead of the server pushing a request down an active stream, it returns resultType: "input_required" along with a description of what it needs answered. The client takes that, gets the answer somehow, and retries the original call with the answers attached under inputResponses. The InputRequiredResult object carries inputRequests, a map of the server's original asks (each one a full elicitation or sampling request, unchanged in shape), plus an opaque requestState value the client has to echo back exactly as received. No memory required on either end between the request and the retry; the state travels with the message instead of living in a socket.
The 2026-07-28 draft spec states the change plainly: the MRTR pattern replaces the previous approach of sending server-initiated requests such as roots/list, sampling/createMessage, or elicitation/create. What survives that swap is the interaction itself, not just the underlying data shape. A tool can still ask for confirmation before deleting something. A server can still ask for a parameter it's missing. An agent can still borrow reasoning from the client side mid-task. What disappears is the requirement that anything stay connected in between those steps, which means any given request can now land on any server instance sitting behind an ordinary round-robin load balancer: no stickiness, no shared state store, no pinned connection.
Worth sitting with that for a second: the philosophy behind sampling, server delegates reasoning, human stays in the loop, isn't the thing that got deprecated. The wire mechanism did. MRTR is a plumbing change dressed up as a spec revision, not a reversal of what sampling was trying to accomplish. For anyone who's already read sampling/createMessage request shapes closely, that knowledge transfers almost directly, since the same elicitation and sampling request objects show up again inside inputRequests, just delivered by a different courier.
What to do with sampling knowledge today given where the spec is headed
For teams already running sampling in production, none of this is a five-alarm fire. The twelve-month deprecation window is real, and there's no emergency migration deadline sitting on the calendar right now. But the direction is unambiguous, and pretending otherwise just trades a planned migration now for a rushed rewrite later.
For teams designing new servers, the spec's own guidance is about as direct as these documents get: evaluate calling the LLM provider's API directly rather than reaching for sampling, and build any mid-call interaction pattern (confirmations, missing-parameter prompts, reasoning handoffs) with MRTR's shape in mind from the start rather than bolting it on after the fact. Given the origin-authentication gaps documented above, that's not just an architecture preference, it's the more defensible security posture too.
For teams building clients, the sampling-support work already invested doesn't evaporate. The concepts, model preference negotiation, human approval checkpoints, tool calling inside a delegated reasoning step, are exactly the concepts MRTR needs implemented too. Learning sampling first isn't wasted effort just because the wire format underneath it is changing. The mechanism aged out, but the idea it was built to express, that a server can borrow a client's brain for a moment without needing one of its own, is sticking around under new plumbing.
Sources
- One Year of MCP: November 2025 Spec Release
- Sampling - Model Context Protocol
- mcp-for-beginners/03-GettingStarted/14-sampling/README.md at main · microsoft/mcp-for-beginners
- Flipping the flow: How MCP sampling lets servers ask the AI for help — WorkOS
- MCPSecBench: A Systematic Security Benchmark and Playground for Testing Model Context Protocols
- MCP Sampling: When Your Tools Need to Think
- stacktr.ee
- modelcontextprotocol.io


