Supply Chain Risks in Third-Party MCP Servers
Compromised servers can manipulate AI agents across many transactions, not just once.

Third-party MCP servers function as a supply chain in the fullest, classical sense of the term: an AI agent delegates trust to external code it did not write, cannot inspect while it runs, and does not control. That delegation is the entire subject of this piece. The Model Context Protocol, introduced by Anthropic, is an open standard that lets AI agents connect to external tools, data sources, and APIs through one interface, which is what makes agentic workflows possible in the first place: an agent can query a database, trigger a job, or act inside an internal application without a developer writing a custom integration for each case.
The architecture behind that convenience has three layers: the host, the client, and the server. The host is the AI application itself. The client is the protocol handler that manages the connection. The server is the bridge to whatever external system the agent needs to reach, and it's where security concerns concentrate, because the server sits between what the AI wants to do and what the organization's systems will actually allow. Every MCP server exposes its capabilities in three forms: tools, which are executable functions; resources, which are structured context the agent can read; and prompts, which are reusable templates. All three are discovered dynamically, at runtime, with no hardcoded integration behind them.
That dynamic discovery is the detail that makes this a supply chain concern: a compromised or malicious server builds a standing dependency, not a one-time integration. A compromised or malicious MCP server can do more than hand over bad data once. It can shape an agent's behavior across many transactions and decisions, continuously, for as long as the agent keeps calling it. A corrupted static library sits still until someone patches it. A corrupted MCP server can manipulate an agent's reasoning in real time, on every call, adjusting what it tells the agent based on what the agent is doing. Antiy CERT confirmed 1,184 malicious skills across ClawHub, the marketplace serving the OpenClaw AI agent framework, a figure that shows what happens once a framework turns a language model from a passive text generator into something that can discover tools, call APIs, read resources, trigger workflows, and act across multiple systems on its own. Once an agent can do all of that, the number of places an attacker can intervene grows with it. Treating each MCP server as a vetted, one-time integration misses what the protocol actually does: it builds a standing, continuously-active dependency between an organization's systems and code someone else maintains.
The protocol's design gaps that make every external server a potential attack surface
MCP was built to solve interoperability, not security, and that priority appears directly in what the protocol does and doesn't handle. It manages the mechanics of connecting an agent to a tool and passing data back and forth. It includes no built-in identity system and no access control. If an organization doesn't bolt on controls from outside the protocol itself, every server an agent trusts inherits whatever permissions it got at setup, and every request that flows to it goes unverified.
Three gaps define the problem. The first is the absence of authorization: an agent inherits whatever permissions its server holds, and there's no per-request check on who is actually asking or whether a given action falls within that requester's scope. The second is unsanitized command execution. Many MCP servers use the STDIO transport, so configuration leads straight to command, and if you don't sanitize inputs from the language model, you can get arbitrary code execution on the server host. Pomerium's 2026 analysis identifies this as a path by which a single malicious prompt can compromise the machine the server runs on, not just the data it handles. The third is the breakdown of the trust model itself. Securing MCP means accounting for a full chain of delegation: who the user is, which client is acting on their behalf, which server is being called, what token authorizes the call, what resource is being touched, and whether the specific tool action requested is authorized in that context. Checkmarx says this chain needs more than API security or OAuth security can give, because neither model was built to track delegation through an autonomous intermediary.
Darktrace's CISO-focused analysis names the resulting condition the "lethal trifecta": access to sensitive data, exposure to untrusted content, and the ability to communicate externally. Any one of those three, on its own, is a manageable risk. A server with access to sensitive data but no path to the outside world is contained. A server exposed to untrusted content but with no access to anything valuable is a nuisance. Put all three in a single server, and the combination becomes something a careful configuration can't undo after the fact. MCP connects AI systems to production infrastructure, so it amplifies risks that already existed in those systems. The properties that make the protocol useful, open discovery, dynamic trust, broad reach across tools and data, are the same properties that make it exploitable. These aren't bugs sitting in a queue waiting for a patch. They describe how the protocol currently works, and every attack mechanism that follows depends on one or more of them.
Tool poisoning and prompt injection: how the agent's reasoning becomes the attack vector
Agents trust tool descriptions and resource content as they encounter them, dynamically, which gives an attacker a way in that never touches any code a security team reviewed: instructions hidden inside the data the agent reads.
Prompt injection works by planting malicious instructions inside ordinary-looking content, a support ticket, a file, a repository issue, so that the agent absorbs them the moment it processes that content as part of its normal job. The GitHub MCP incident shows the mechanism concretely. A server had access to both public and private repositories. A user asked the agent to review repository issues, nothing out of the ordinary. An attacker had planted a prompt-injection payload inside a public issue. The agent read it, followed it, and then put sensitive private data into a pull request it generated, so that data went out without the user ever intending it. That single sequence combined all three legs of the lethal trifecta: access to private data, exposure to a malicious instruction, and a channel to push the result out. Checkmarx's 2026 analysis names the underlying mechanism: when the trust boundaries between user, host, client, server, and downstream systems blur together, a document the agent merely retrieves can push the model into calling a tool the user never actually approved.
Tool poisoning works differently but depends on the same blind trust. A server operator, or an attacker who has compromised one, can control the metadata attached to a tool: its description, its parameters, the instructions embedded in how it presents itself to the model. That operator can update the hidden instructions in that metadata, but the name and summary a human reviewer sees stay unchanged. The agent reads the metadata as authoritative and acts on instructions no person ever saw. The OWASP MCP Top 10 treats this as serious enough to name separately, listing Tool Poisoning as MCP03:2025 and Context Injection and Over-Sharing as MCP10:2025, distinct line items in a formal taxonomy of how MCP deployments fail.
The same logic extends across servers. A malicious server doesn't need to attack its target directly. It can embed instructions inside third-party resources that its own tools touch, so when a user invokes that tool, the instructions chain outward into calls against other, trusted MCP servers the agent happens to have access to. The damage isn't contained to the server that was compromised. It reaches every server the agent can reach. Darktrace calls this cross-agent contamination: in multi-agent environments, shared servers or shared context stores let malicious context spread from one agent to another, turning a single compromised source into a systemic problem across an entire deployment.
The confused deputy problem and over-permissioned access
A second class of exploitation works without injecting anything into the system. It works by using a server's own legitimate authority against the system that authority was meant to protect. Because MCP lacks strict authentication, a server will act on a request without verifying who is actually behind it, so an attacker who can't steal the server's permissions directly can still get the server to exercise them on the attacker's behalf.
If an agent holds broad write permissions, you can prompt it to modify or delete critical resources in response to a request that looks legitimate but comes from a lower-privileged agent or user somewhere upstream. The server has the permissions. The attacker only has to supply a convincing request. Execution follows without a challenge at any point in between. Darktrace's April 2026 briefing flags this as a particular concern for MCP deployments: agents interact across tools and services autonomously, and that autonomy makes privilege misuse much harder to catch than a stolen credential, which at least leaves a login event you can find.
Over-permissioned access compounds the exposure before any attacker shows up. MCP servers often hold broad standing access to file systems, APIs, and databases, and that access persists whether or not the agent is using it at a given moment. Agents can also chain multiple tools together into sequences of action that no human operator ever signed off on step by step. Darktrace identifies this tool-chaining behavior, under agents that already hold more privilege than their task requires, as one of the clearest ways MCP amplifies risk that already existed before MCP entered the picture. An overly permissive agent can modify database configurations or take other damaging actions, and no single step in that sequence looks obviously malicious on its own. The FINOS AI Governance Framework names privilege escalation as a technical compromise scenario, and it notes that a compromised MCP server can reach sensitive data well beyond what any approved use case called for. That outcome is what the current trust model produces by default, a point the Cloud Security Alliance's research (commissioned by Zenity) backs with a specific figure: between 47 and 53 percent of organizations report that an AI agent has exceeded its permissions or caused an incident. The confused deputy problem doesn't need an attacker to actively inject anything into the system. An agent behaving exactly as designed, calling the tools it was built to call, is still exploitable whenever its permission scope is wider than its task.
Rug pulls, typosquatting, and the instability of third-party trust over time
The most underappreciated risk in the MCP supply chain is time. A server approved during deployment review is not guaranteed to be the same server running six months on, and the protocol has no mechanism built in to detect or stop a behavioral change that happens after approval.
A rug pull is exactly that: a third-party server that looked safe during review quietly starts behaving differently later, with no code change visible to the organization that approved it in the first place. If the server's own operators push updates and patches, those can introduce backdoors, corrupt data, or manipulate logic, and the operators may never know it happened. FINOS's governance framework gives this its own name, MCP Server Update Poisoning, and treats it as separate from a server being compromised at the point of initial deployment. The risk persists over time, too. FINOS notes that a compromise of this kind can persist undetected for extended stretches of time, quietly affecting many transactions and decisions before anyone notices, a pattern that lines up with supply chain compromises making up a growing share of significant cloud incidents recently.
Typosquatting exploits the same mechanism that makes MCP convenient to use. Developers and agents select servers by name, and a server registered under a name close enough to a well-known, legitimate one can intercept traffic meant for the real thing. Pomerium's 2026 analysis groups typosquatting with rug pulls as compounding risks, since both take advantage of the same missing piece: MCP has no cryptographic way to verify a server's identity.
The SANDWORM_MODE campaign, which surfaced on February 20, 2026, shows both vectors working together in one operation. At least 19 malicious npm packages impersonated popular developer utilities and AI coding tools through typosquatted names, close enough to the real packages that they could catch developers who were moving quickly. Behind that disguise, the rogue server advertised three tools that sounded entirely routine: index_project, lint_check, and scan_dependencies. Each one was built to harvest SSH keys, cloud credentials, npm tokens, and other environment secrets from developer machines and CI pipelines. Nothing about the names would have raised a reviewer's attention. FINOS's framework adds one more dimension to this instability: insider threats to MCP infrastructure, where someone with legitimate access to a server's backend deliberately introduces a backdoor or corrupts data from inside. Rug pulls and typosquatting show that outside attackers aren't the only way this trust problem gets in.
What happens when the SDK itself is the vulnerability
Supply chain risk in MCP doesn't stop at the server an organization chooses to deploy. The SDKs that implement the protocol sit underneath every one of those servers, and a flaw in an SDK travels into every server built on top of it, regardless of how carefully that specific server was reviewed.
CVE-2026-25536, carrying a CVSS score of 7.1, is a case in the TypeScript SDK that shows how far that reach extends. The flaw caused a single server instance to reuse one McpServer or Server object, along with its transport instance, across multiple clients at once. Responses meant for one client got routed to another, leaking one user's data straight into a different user's session. It didn't matter which organization happened to run the affected code. Every server built on those SDK versions carried the same flaw, whether or not the team running it had any idea the SDK itself was the weak point. That is the shape of the MCP supply chain: a server an organization vetted, built on a protocol with no authentication built in, running code that depends on an SDK that organization never inspected.
Sources
- MCP Server Supply Chain Compromise
- MCP Security Crisis: Systemic Design Flaws in AI Agent Infrastructure
- Model Context Protocol Threat Modeling and Analyzing Vulnerabilities to Prompt Injection with Tool Poisoning
- Parasites in the Toolchain: A Large-Scale Analysis of Attacks on the MCP Ecosystem
- MCP Tool Poisoning: Adversarial Hijacking of AI Agent Workflows


