The Context Window

MCP Server Authentication in Production APIs

OAuth 2.1 becomes mandatory as MCP matures past its early prototype phase.

Editor at Large · · 12 min read
Cover illustration for “MCP Server Authentication in Production APIs”
Model Context Protocol · September 5, 2026 · 12 min read · 2,746 words

MCP (Model Context Protocol) went from a November 2024 launch to a fixture in enterprise AI platforms in about eighteen months, and its authentication spec spent most of that time playing catch-up. Teams built production servers against guidance that kept moving under their feet, and a lot of what shipped in year one now reads like an unfinished sentence someone forgot to close. What follows walks through the layers a production MCP deployment runs on today, transport choice, the OAuth 2.1 flows now required, token scoping, the discovery chain that connects clients to authorization servers, and the exact spots where that stack still breaks.

None of this is a knock on the engineers who wired up early servers with API keys and good intentions. The correct approach genuinely didn't exist yet. A good chunk of the public MCP surface has drifted toward a state where security hasn't kept pace with maturity elsewhere in the spec, and that gap deserves more attention than it's getting.

How the spec's treatment of authentication changed across versions

Four releases matter here. At launch in November 2024, MCP shipped with stdio and HTTP+SSE as transports, and the authentication guidance amounted to a shrug: do something reasonable, details to follow. The details arrived in June 2025, when the spec reclassified MCP servers as OAuth Resource Servers, made RFC 9728 discovery mandatory, killed the fallback default endpoints earlier drafts allowed, and deprecated SSE transport outright.

November 2025 is the version most production teams should be building against right now. It made OAuth 2.1 the standard for remote servers, made PKCE mandatory instead of merely recommended, and introduced step-up authorization for requests that need more scope than the current token grants. A July 2026 release candidate goes further: it hardens issuer validation using RFC 9207, binds client credentials to the issuer that actually issued them, and deprecates Dynamic Client Registration in favor of client metadata documents (more on why DCR earned that fate a few sections down).

Here's the trap that keeps catching people: a blog post or vendor doc written before June 2025 describes a version of the protocol whose server role and discovery mechanism have since changed underneath it. Follow pre-June guidance today and a client ends up pointed at endpoints that no longer exist. Google Cloud's remote MCP servers implement the November 2025 requirements, and they're a decent reference for what a conformant deployment looks like at real scale.

Diagram: MCP Auth Spec: Four Releases, Four Shifts. Visualizes: Show the evolution of MCP authentication requirements across four named releases on a timeline.

Transport choice determines whether authentication is even possible

MCP supports two transports, and only one of them can be secured in any meaningful sense, which is worth being blunt about.

STDIO is the local option: client and server talk over standard input and output, usually wired through environment variables, with zero authentication because trust comes from the host process boundary itself. That's fine for one developer on one laptop, but once an organization has dozens of developers each running local MCP clients against a handful of servers apiece, that boundary stops meaning anything. There's no shared control plane, no central audit trail, and no way to rotate a credential without touching every host by hand. If something goes wrong, someone has to manually reconstruct which process on which machine did what, after the fact, which is a miserable way to spend an afternoon.

Streamable HTTP, introduced in March 2025 and carried into the November 2025 spec, replaced the older HTTP+SSE transport (deprecated June 2025; don't build anything new on it). It's a single-endpoint design that supports authentication, horizontal scaling, stateless deployment, and resumable streams, and it's the only transport that gets a deployment to centralized auth, audit logging, or policy enforcement.

Picture the team that starts on stdio because it's faster to prototype against, with a plan to "add auth later." Later arrives, and it turns out the transports aren't interchangeable; swapping one for the other is closer to a rewrite than a config change. The shortcut in month one turns into the expensive lesson in month six.

What OAuth 2.1 actually requires of MCP implementations

OAuth 2.1 is now mandatory for HTTP transport, and each requirement maps to a specific failure MCP kept running into before the spec closed the gap.

PKCE (Proof Key for Code Exchange) is mandatory, with S256 as the required challenge method wherever it's feasible. This closes the authorization code interception window, which matters more for agents running headless flows with no browser in the loop than it does for a normal web login. Implicit grant is gone entirely, and good riddance: early agent frameworks liked it because it was simpler to wire up, but it puts access tokens straight into redirect URI fragments, browser history, and referrer headers. In an enterprise environment with logging turned on everywhere, that's a token sitting in a log file somewhere it was never supposed to land.

Resource Indicators, from RFC 8707, are required too. A client has to bind a token explicitly to the specific MCP server it's calling, so a token minted for Server A can't get replayed against Server B. Servers themselves act as resource servers: they check tokens issued by somebody else, and they don't hand tokens out themselves. That distinction sounds like paperwork until it clicks that it's the same distinction that keeps the confused deputy problem, covered further down, from happening in the first place.

Step-up authorization is new as of November 2025, and it solves a real annoyance in multi-capability agents. If a request needs scopes the current token doesn't have, the server returns a 403 with a header listing what's missing, and the client re-runs authorization to pick up the extra scope. It's progressive consent, gathered as needed rather than front-loaded before an agent might ever use it. Pair that with short token lifetimes, an hour is the going recommendation, and the blast radius of a stolen token shrinks to something a security team can actually live with.

How the discovery chain connects clients to the right authorization server

Diagram: The Discovery Chain: Five Steps from 401 to Token. Visualizes: Illustrate the five-step OAuth discovery sequence the November 2025 MCP spec requires before a client can authenticate.

Once the June 2025 spec killed the fallback default endpoints, clients needed a real way to find the correct authorization server instead of guessing at /authorize and /token. Any implementation still hardcoding those paths is running against a version of MCP that doesn't exist anymore.

Here's how it actually runs. A client hits the MCP server with no token, gets a 401 back, and that 401 carries a WWW-Authenticate header pointing to a resource_metadata URL. The client fetches that URL, which is the Protected Resource Metadata document defined in RFC 9728 (living at /.well-known/oauth-protected-resource), and that document names the authorization server in its authorization_servers field. The client then fetches that authorization server's own metadata, defined in RFC 8414 at /.well-known/oauth-authorization-server, to learn what endpoints and grant types it actually supports. Only then does the OAuth 2.1 flow start.

Clean design on paper, but in practice, only a small number of OAuth providers have implemented the full set of features the spec now assumes, and most existing providers would need real engineering work to catch up. That gap between what the spec requires and what a given provider actually ships is the single most common source of integration breakage. Worth naming a specific case: an Atlassian MCP Server GitHub issue from April 2026 describes clients following the RFC 9728 chain correctly and still failing to locate Atlassian's authorization server, which breaks Dynamic Client Registration and drops the connection. The client code wasn't wrong, and the spec wasn't wrong either. The provider just hadn't caught up yet, and that's worth remembering the next time a "compliant" integration falls over for no obvious reason.

Why Dynamic Client Registration is a security liability at enterprise scale

Dynamic Client Registration (DCR) lets a client register itself against an endpoint automatically, no manual setup required. Reasonable design, in theory, for architectures where clients come and go constantly and nobody wants to pre-register each one. Less reasonable once you actually watch what an open registration endpoint invites, and at enterprise scale it should be treated as a liability, not a convenience.

An anonymous, self-service registration endpoint is hard to monitor and harder to audit; it's an attack surface built for automated abuse, since nothing stops a script from registering clients at scale continuously. Anonymous registration also means there's no verified organizational identity behind a given client ID, which makes IP reputation checks or risk-based access controls close to worthless. The July 2026 release candidate deprecates DCR in favor of client metadata documents, and that's not a cosmetic tweak; it's the spec community watching how DCR behaved once it left the whiteboard and deciding the tradeoff wasn't worth it.

The production alternative is unglamorous: register clients by hand, with verified identities, once per environment, then deploy that registration everywhere it's needed. Less dynamic, a lot more auditable. There's a sharper reason to care beyond hygiene, too: the confused deputy attack described below gets meaningfully easier to pull off when client IDs are dynamically registered and not cryptographically tied to a specific issuer.

Token passthrough and the confused deputy: the two anti-patterns the spec explicitly forbids

The spec doesn't hedge on this one. An MCP server must not accept a token that wasn't issued to it, and when that server needs to call an upstream service on the user's behalf, it has to get a separate token scoped specifically to that upstream service. Passing the original token through is forbidden, full stop, no exceptions carved out for convenience.

Why does passthrough matter so much? Once it's allowed, audience checks stop meaning anything, the downstream API's own rate limiting and validation get bypassed entirely, and the audit trail ends up recording the wrong actor as the source of the request. Stack the confused deputy pattern on top: a downstream API sees a forwarded token and treats it as though the MCP server itself issued and validated it, granting access the actual user was never entitled to. The server becomes an unwitting middleman handing out permissions it never had the authority to hand out in the first place.

One documented variant is worth walking through in detail: a proxy MCP server using a static client ID against a third-party authorization server, combined with dynamic client ID registration, opens a door for an attacker to craft a link that reuses an existing consent cookie, skip the consent screen entirely, and redirect the resulting authorization code to a URL the attacker controls. The fix the spec requires has two parts, and both have to be present: per-client consent enforced before the third-party flow even starts, and exact-match redirect_uri validation with zero tolerance for wildcards. Passthrough and this authorization-code variant trace back to the same root mistake, treating the MCP server as a transparent pipe instead of a bounded resource server with its own identity and its own audience-scoped tokens.

Credential types in use and where each one fits

Three credential types show up across production MCP deployments, and they are not interchangeable, whatever the convenience of pretending otherwise might suggest.

OAuth 2.1 access tokens are the recommended shape for good reason: short-lived, bound to a specific audience, limited in scope. The entire threat model the spec is built around assumes this credential sits at the center. API keys are simple to issue and use, which is exactly why a large share of early MCP servers picked them, and why a large share of public servers today are still running on them. The catch is that API keys don't expire on their own, don't support fine-grained per-user permissions, and grant service-level access instead of user-level access. Once one leaks, it stays useful to whoever has it until somebody manually rotates it. That's a workable tradeoff for a tightly controlled internal tool with an enforced rotation schedule. It is not a workable primary credential for anything multi-user or reachable from outside the network, and treating it as one is how a security review turns into a bad afternoon.

JWTs sit in between. They can carry identity claims directly, user, role, tenant, which is handy for passing context into a server without extra lookups. That convenience comes with a catch of its own: the server has to actually validate the signature, check the claims, and enforce authorization logic correctly, every single time. A JWT that nobody bothers to check for the right audience, issuer, and expiration carries the same risk as a plain string, just dressed up with more fields. The real guidance is matching the credential type to the actual threat model the server faces, not grabbing whichever one was fastest to wire up on a Tuesday afternoon.

How enterprise IdP integration fits into the MCP auth architecture

The spec's core separation, the identity provider issues tokens and the MCP server only verifies them, is what keeps this whole architecture from turning into a pile of bespoke auth code bolted onto every server individually. That separation is arguably the single most useful design decision to come out of the post-June-2025 spec.

Organizations already running OIDC-based identity systems can register MCP servers as clients with the IAM provider they already run. No parallel auth stack, no second identity system to keep in sync with the first. A handful of practices make this work well in production. Use token exchange under RFC 8693 instead of passing through the user's original OAuth token, so every downstream action attributes correctly to the right principal and the audit log stays clean. Cut scopes down aggressively, since least privilege is cheap to design for up front and expensive to retrofit later. Keep tokens short-lived and pair them with proof-of-possession under RFC 9449 (DPoP) to block replay attacks, and where a token alone doesn't carry enough context for a fine-grained decision, Rich Authorization Requests under RFC 9396 fill that gap per request.

The payoff is straightforward: once an MCP server looks, from the IdP's point of view, like just another OAuth resource server, centralized policy enforcement, audit logging, and credential rotation all apply to it automatically, the same way they already apply to the rest of an organization's API surface. No special case, no separate playbook to maintain.

What the current vulnerability landscape reveals about where teams are actually failing

Here's the uncomfortable part. Despite the spec maturing considerably by November 2025, a substantial share of publicly reachable MCP servers still run with no authentication at all. The two sections above explain most of why: stdio deployments that quietly escaped their intended network perimeter, and API-key servers where nobody ever enforced a rotation policy.

The vulnerability categories security audits keep turning up map almost exactly onto the failure modes already covered here. Missing or trivially bypassed authentication is what happens downstream of picking the wrong transport or the wrong credential type. Token passthrough and audience confusion are what happens when a server gets treated as a transparent proxy instead of a bounded resource server. Path traversal and server-side request forgery show up when a server acts on file paths or URLs without checking what the calling principal is actually entitled to touch, which is a scope enforcement gap wearing an input-validation costume. Tool poisoning is its own animal entirely: it's an attack on content, not identity, and no amount of correct OAuth stops it, because it requires verifying the source of a tool's definition, a problem authentication was never built to solve.

One CVE deserves naming on its own. CVE-2025-6514, affecting mcp-remote, scored 9.6 on CVSS for OS command injection, and correct OAuth on the front door doesn't help much if the server mishandles what comes through it once the door opens. In multi-agent setups where several MCP servers share a trust boundary, a compromise on one server can cascade into the others sharing that boundary, which is exactly why strict per-server audience scoping matters as much as it does: it's the control that actually limits how far a single breach travels.

So where does that leave a production team today? The spec itself now offers a coherent, workable security model, arguably the first version of MCP where that's been true. What's left is implementation fidelity and plain operational discipline: short token lifetimes, client registration that's actually verified, audience validation enforced without exception, and audit logging centralized instead of scattered across a dozen hosts. None of that needs new infrastructure. It needs every MCP server, from the first line of code, treated as a bounded OAuth resource server with its own identity, built with authentication from the start rather than bolted on after the first incident.

Sources

  1. stackoverflow.blog
  2. truefoundry.com
  3. boomi.com
  4. docs.cloud.google.com
  5. practical-devsecops.com
  6. arxiv.org

More in Model Context Protocol