The Context Window

Least Privilege Design for AI Agent Tool Permissions

AI agents need permission controls designed for rapid, autonomous action, not static human roles.

Editorial team · · 11 min read
Cover illustration for “Least Privilege Design for AI Agent Tool Permissions”
LLM Security and Trust · October 5, 2026 · 11 min read · 2,376 words

Least privilege for AI agents cannot be built on the same foundation as human identity and access management, because the two systems make incompatible assumptions about how privilege moves. Traditional IAM assigns permissions to human identities tied to stable job roles: an employee gets a title, the title maps to a set of permissions, and those permissions stay fixed until the employee's job changes. An AI agent acts autonomously across systems in sequences that no single human approves at each step, and that fact alone dissolves the premise that makes static role assignment safe.

The difference shows most clearly in how privilege drift happens. In legacy IAM, drift is slow. It follows onboarding, offboarding, and job changes, cycles that unfold over weeks or months and give security teams time to review what access has accumulated. Agent-driven drift moves much faster. Development teams over-provision OAuth scopes so a workflow doesn't break, service accounts get reused across deployments without anyone re-examining what they touch, and permissions stack on top of each other with no one tracking what the combination adds up to.

The danger compounds. An agent with access to email, a file store, a ticketing system, and a code repository can look reasonable when each integration is reviewed on its own. Microsoft's July 2026 guidance points out that the combination lets the agent correlate data across all four systems and take actions that no one explicitly authorized as a whole: the sum of four modest grants can produce an effective permission set broader than any single reviewer evaluated.

The actor that chooses an action, the user who requested the task, the runtime holding the credentials, and the service receiving the request can all be different entities, which is a harder problem than the compounding risk alone. The paper "Authorization Architectures for Tool-Using AI Agents" identifies this as the confused deputy problem, the clearest lens available for understanding agentic authorization failures. A privileged agent gets tricked into misusing its own authority because it has no reliable way to tell which of its several authorities should govern the action in front of it. Agent safety cannot be reduced to model alignment. The actions agents take need explicit authorization controls that hold regardless of what the model "intends."

Where permission failure originates in an agent's execution chain

Permission failure in an agentic system doesn't start at the moment an account gets provisioned. It can originate at any of several distinct points along the execution chain, and each point is a separate surface that needs its own scrutiny.

Five layers deserve independent treatment. The first is tool inventory: which tools the agent can even see as options. The second is invocation: whether the step the agent is on actually calls for the tool it just selected. The third is argument: the specific object, recipient, query, file path, dollar amount, or payload the agent passes into that tool. The fourth is data: how much of the tool's output gets returned to the model, and whether sensitive fields inside that output are redacted before the model ever reads them. The fifth is delegation: whether a sub-agent ends up holding more authority than the subtask in front of it requires.

Delegation is most likely to go wrong quietly. When an agent hands work to a sub-agent, many deployments issue fresh credentials at that handoff rather than passing along a bounded version of the original grant. A downstream agent can end up operating with permissions the original user never authorized. "Authorization Architectures for Tool-Using AI Agents" describes this as the chain of trust breaking across multi-hop delegation, and the break is structural: nothing in a typical credential-issuance step checks that the new grant stays inside the bounds of the old one.

The scale of the tool surface itself is growing fast. A study of public Model Context Protocol repositories counted 177,436 agent tools created between November 2024 and February 2026, and an increasing share of them are action-oriented. The boundary where an LLM's internal plan turns into an action visible outside the model, a sent email, a deleted file, a payment, is expanding every month that count runs.

Temporary access adds a failure mode of its own: a grant issued for one task, with no expiry mechanism attached, becomes permanent access simply because nobody revokes it. Microsoft documents a related pattern directly. A team provisions a "Reader" role because the first use case only needs to read data. The workflow then expands to include fixing the issues it finds, and rather than redesigning the role to match the new task, the team just grants something broader and moves on. The expansion is incremental, easy to justify in the moment, and rarely revisited once the workflow stabilizes.

Exploitation of These Failure Points in Production Systems

These failure points are not hypothetical. Overprivileged agents have already been exploited in production, through routes ranging from prompt injection to stolen credentials, with consequences that include data exfiltration, compromised supply chains, and direct financial loss.

The clearest documented case is EchoLeak, tracked as CVE-2025-32711 and patched by Microsoft in June 2025 with a CVSS score of 9.3. Researchers at Aim Security disclosed it as the first documented zero-click prompt injection exploit against a production AI system. A single crafted email, one that required no interaction from the victim beyond ordinary use of Microsoft 365 Copilot, caused Copilot to pull internal files and send their contents to a server the attacker controlled. No one clicked a link. No one opened an attachment. The exploit worked because Copilot's own permission scope let it reach the files in the first place, and the injected instruction simply redirected what the agent already had the authority to do.

A common thread runs through incidents like this one: when an attacker steals an agent's OAuth token or API key, the attacker inherits the full effective authority that agent holds, often spanning several SaaS applications at once. The theft itself doesn't need to be sophisticated. The payoff is large because the agent's own permission footprint was large to begin with.

A subtler failure occurs even where human oversight is built into the workflow. Research on user-authored permission policies published in August 2026 found that when users are asked to approve agent actions at runtime, a large share of overreach actions still go through after the user approves them. The gap sits between what a person says they want in the abstract and what they actually approve when a specific request appears in front of them. Human-in-the-loop controls leave a gap between approval and actual protection.

The Principal Hierarchy as the Right Unit of Authorization Design

The moment a tool gets invoked is where enforcement has to happen, but that moment only produces a sound decision if it can be traced back through an unbroken chain of delegated authority to the human who set the task in motion. Treating the tool call as the unit of design misses this entirely, because a tool call looks identical whether it is authorized or not. What distinguishes the two is the chain behind it.

"Authorization Architectures for Tool-Using AI Agents" proposes organizing the whole design problem around a principal hierarchy: human user, then operator or deployer, then orchestrator agent, then sub-agent, then tool endpoint. Every consequential action needs to trace back to a human principal, stay bounded by what that human actually delegated, and remain contestable after the fact. Few documented deployments satisfy all three properties reliably end to end, which is precisely what makes the hierarchy a design target rather than a solved problem.

Microsoft's guidance translates this into a concrete practice: separate duties. When a workflow includes both gathering evidence and acting on it, read and write should sit behind different roles or different tools entirely, and high-impact actions, deleting something, exporting data, changing a privilege, should require a step-up approval rather than riding on whatever permission got the agent that far.

Splitting strategic planning from tactical execution creates a natural checkpoint: a Verifier component inspects what the Planner proposed before the Executor is allowed to act on it, confirming the proposed steps meet security requirements before anything reaches an external system.

Microsoft also identifies the question that incident investigations run into most often, and it's an identity question rather than a technical one: is the agent acting under its own identity, under a scope delegated by a user, or under some blend of the two that no one defined clearly in advance? That ambiguity decides who is accountable for what the agent did and what approvals should have been in place before it did it, and it becomes visible in the middle of an incident, when the answer is needed immediately and isn't available. Microsoft recommends treating every agent as a first-class principal in its own right. Give it a lifecycle-managed identity, assign it explicit roles, scope its permissions tightly, and limit its tool usage to a preconfigured manifest.

How task-context-driven permission scoping works in practice

Granting permission based on task context rather than a fixed role requires a mechanism that can figure out, at the moment a task starts, what that specific task actually needs, then hold the agent to that scope at the argument and data layers as well as at the level of which tools it can see.

Progent, introduced in April 2025 and updated in May 2026, works this way. A large language model generates an initial privilege policy directly from the user's task description, then updates that policy as new information comes in during execution. Every update gets checked by a deterministic SMT solver before it takes effect. Narrowing the policy happens automatically. Expanding it requires explicit approval. That asymmetry produces a property the researchers call monotonic confinement: the agent's effective action space can only shrink without approval, even in a scenario where the LLM generating updates has been manipulated by adversarial input.

Intent-Governed Access Control, a 2026 framework, approaches the same problem from a different angle. It derives a short-lived certificate from a request it trusts, narrows the agent's authorized tool manifest down to what that certificate covers, and checks the effects of a proposed tool call and its payload before letting the call through.

Scoping the tool manifest is not sufficient on its own. Even when a specific tool invocation is authorized, the payload passed into it, the scope of a query, or the volume of data returned to the model each need separate enforcement. A tool can be the right tool for the job and still return far more than the task requires unless sensitive fields are stripped out before the model sees them. Argument-level and data-level control are a different layer of the problem than deciding which tools an agent is allowed to call, and both frameworks treat them as such.

The unresolved tension between dynamic policy generation and deterministic enforcement

Using a language model to generate permission policy introduces the same unpredictability into the enforcement layer that the enforcement layer exists to guard against. No framework available today has fully closed that gap, though some hybrid designs narrow it meaningfully.

The dilemma has two bad options at its poles. A manually written policy needs security expertise to build and cannot anticipate task contexts nobody thought of in advance. An LLM-generated policy adapts to whatever context actually shows up but carries none of the rigor a security team would want to rely on. Progent's SMT-solver hybrid tries to split the difference by giving the LLM a narrow job: propose narrowings. A deterministic checker sits behind it and blocks any expansion that isn't explicitly approved.

The underlying concern has not gone away. If the policy-generating LLM is manipulated into proposing something broader than it should, the SMT check is what catches the expansion attempt. But the quality of the starting policy, the one the LLM derives from the task description before any manipulation enters the picture, still depends entirely on how well that model understood the task in the first place. A solver can block an illegitimate expansion. It cannot tell you that the initial policy was too generous to begin with.

Research on user-authored permission policies found that participants chose "ask" for a large share of their standing rules, which just pushes most of the actual decisions back to the moment the agent is already running. A standing policy built to preserve case-by-case judgment ends up protecting less than it looks like it should, because people approve overreach actions when those actions are presented individually, one at a time, outside the context where the standing rule was supposed to apply.

The strongest concrete guarantee documented in this space remains monotonic confinement, the property that an agent's action space can only narrow without explicit sign-off. Its strength is still conditional on the correctness of the policy it started from. Permissions narrow enough to be safe might also be narrow enough to stop an agent from finishing a novel multi-step task without constant interruption, which cuts against the reason agentic systems are useful.

Regulatory Frameworks, the OWASP Agentic Taxonomy, and These Design Principles as Requirements

Regulatory bodies and standards organizations in multiple jurisdictions are moving task-scoped, dynamically revocable permission design out of the research literature and into the language of formal requirements. What began as an architectural recommendation in papers like "Authorization Architectures for Tool-Using AI Agents" is becoming the kind of thing an auditor checks for rather than the kind of thing only a security researcher recommends.

The direction of that shift matters as much as its existence. A principal hierarchy that traces every action to a human, bounds it to what was delegated, and leaves it contestable afterward is not just good design practice anymore. It is becoming the kind of property regulators expect to see demonstrated, documented, and tested, which changes the incentive for organizations deploying agents from "this reduces our risk" to "this is what we will be asked to show." The incidents already on record, the scale of tool creation already measured, and the gaps already documented between user approval and actual protection all point in the same direction: the design principles this piece has laid out are no longer optional hardening. They are becoming the baseline a production agent is expected to meet.

Sources

  1. Least privilege for AI agents: Identity, access, and tool binding
  2. Intent-Governed Tool Authorization for AI Agents Genliang Zhu Accentrust
  3. Authorization Architectures for Tool-Using AI Agents
  4. Do User-Authored Permission Policies Improve Protection Against AI Agent Overreach?

More in LLM Security and Trust