The Context Window

ReAct vs Plan-and-Execute Agent Patterns

ReAct adapts step-by-step while Plan-and-Execute locks in the full sequence upfront.

Staff Writer · · 13 min read
Cover illustration for “ReAct vs Plan-and-Execute Agent Patterns”
AI Agent Architectures · September 13, 2026 · 13 min read · 2,887 words

ReAct and Plan-and-Execute solve the same problem in opposite directions, and most teams pick between them by accident rather than by design. ReAct reasons one step at a time, folding each new piece of information into its next move. Plan-and-Execute writes the whole sequence of steps before acting on any of them. Per the LangChain State of AI Agent Engineering Report (2026), 57% of AI teams now run agent patterns in production, so the choice between these two loop structures has stopped being academic, and getting it wrong is an expensive way to find out which one you needed. Most teams still say "build an agent" without naming which loop they mean, and that gap, more than any model limitation, is where a lot of production trouble starts.

Neither pattern is a matter of model capability. The same underlying LLM can run either loop; what changes is how the steps get wired together and when the plan, if there is one, gets written down. This piece covers single-agent reasoning patterns, not multi-agent orchestration. LangGraph, CrewAI, and AutoGen are the plumbing. ReAct and Plan-and-Execute are the architecture running through the pipes, and conflating the two is the first mistake worth ruling out.

How ReAct structures reasoning as a repeating loop of thought, action, and observation

ReAct comes from Yao et al.'s 2022 paper "ReAct: Synergizing Reasoning and Acting in Language Models" (arXiv:2210.03629). The structure is simple to state even though its consequences aren't: the agent runs Thought, Action, Observation, Thought, Action, Observation, on and on, until it decides it has enough to answer. Each Observation, whatever a tool call or search or database query returns, feeds straight into the next Thought. The agent reasons about what it just learned, never about what it guessed at the start.

No global plan sits anywhere in this loop. Ask the agent what its plan is at any given moment, and the honest answer is that it doesn't have one beyond its next intended action. That sounds like a weakness, and it is one, but it's also the source of ReAct's real strength: when a tool fails or returns something odd, the very next Thought absorbs that fact and the agent changes course without any separate re-planning step bolted on. Adaptation is built into the loop itself, not added after the fact.

The cost of that same design is nearsightedness. The agent only ever sees one step ahead, so it can't optimize a path it never lays out in full. On the ALFWorld benchmark, ReAct beat named imitation-learning and reinforcement-learning baselines by 34 percentage points in absolute success rate. That's a striking number, though it's tied to that specific benchmark and shouldn't be read as a guarantee that transfers to arbitrary production tasks.

ReAct should be the default, full stop, and teams should only reach for something heavier once a task actively proves it needs one. LangGraph's create_agent function (previously create_react_agent, now deprecated as a standalone name but still the underlying behavior), LangChain more broadly, and NeMo all ship a ReAct loop as the out-of-box starting point. When someone says "agent" with no further qualifier, this loop is almost always what they mean. It's also easy to audit: every Thought gets logged as a discrete step, so when a run goes wrong, the trace shows exactly where the reasoning turned. That matters in regulated industries, where an auditor wants the chain of reasoning, not just the final output.

How Plan-and-Execute separates the act of planning from the act of doing

Plan-and-Execute splits the job into three roles. A Planner writes the full ordered sequence of steps up front. An Executor works through each step, and because each step is narrow and well-defined, the Executor can run on a cheaper or smaller model, sometimes even deterministic code with no LLM involved at all. A Re-planner gets triggered only when something during execution breaks the original plan's assumptions.

That division of labor is the whole point. The Planner does all the strategic thinking once, in one expensive call. The Executor never reasons about the big picture; it just completes one sub-task at a time, which means it can run lighter without losing accuracy on the piece it's actually assigned. Some implementations add a Verifier component, described in Del Rosario et al. (2025, arXiv:2509.08646), affiliated with SAP and the University of Oregon, which checks the plan for logical soundness and security compliance before execution begins. That step earns its keep most when running the wrong plan would cause damage nobody can undo.

Sophisticated versions nest ReAct inside the Executor. Plan-and-Execute sets the mission at the strategic level, and each individual step then runs its own adaptive ReAct loop to handle whatever nuance shows up during that one step. That's a native hybrid, not a compromise, and it shows up in some of the more advanced deep research systems in production today.

One underrated benefit: the plan is an artifact you can look at before a single token gets spent on execution. A team can review the proposed steps, catch an obviously wrong approach, and stop it before it runs. LangChain's original 2023 reference implementation generated one plan at the start and never revisited it, which is brittle by design and has no business being anyone's production setup in 2026. Newer implementations, and LangGraph has released agent architectures built around this style that add conditional re-plan gates, designed to address limitations of purely reactive loops. Re-planning doesn't happen automatically, though. It's a mechanism someone has to design in on purpose; skip it, and Plan-and-Execute inherits all the fragility of a single frozen plan.

Where ReAct's step-by-step adaptability becomes a liability

The same property that makes ReAct flexible fails in three specific, related ways.

Repeated reasoning is the first. Without a fixed plan anchoring the sequence, the agent can re-examine a decision it already made, spending tokens on reconsideration instead of forward progress. Goal drift is the second: each individual reasoning-action pair can look locally sensible while the overall trajectory quietly wanders off the objective, since nothing in the loop ever checks whether the full sequence still holds together. Faulty grounding is the third. Stale or contradictory information lands in an Observation, gets trusted in the next Thought, and then gets cited again downstream. Practitioners call this context poisoning: one bad piece of information enters the loop and keeps getting repeated as if it were true.

A more specific version shows up often enough to deserve its own name: verifier stall. The agent calls a verify_result tool, gets a failure, rewords its arguments slightly, and calls it again, with no memory of the earlier failed attempt once that attempt scrolls out of the context window. It burns tokens in circles, and it does so quietly, since each individual call looks reasonable in isolation. The fix requires more than a global step cap alone. It's a cap on repeated calls to that specific tool, tracked by name, not just a ceiling on total iterations.

Cost compounds this. Every iteration re-sends the accumulated history to the model, so billed input tokens can grow quadratically with the number of iterations, not linearly. A ten-step task, in other words, can cost far more than ten times what a single-step task costs. Without a hard iteration cap and a timeout, the loop just keeps running until the context window fills or the budget runs out, and those two settings aren't optional extras. They're load-bearing configuration, and skipping them is how a team ends up staring at a surprise bill.

From the outside, all three failure modes look identical: a long trace that never converges. Diagnosing which one actually happened, repeated reasoning, drift, or bad grounding, usually takes more than reading the log.

Where Plan-and-Execute's upfront commitment creates its own risks

The Planner commits to a full sequence before it has seen a single tool output. If step 2 returns something the Planner never anticipated, step 3 is already written, and it's already wrong. Every remaining step can then execute flawlessly against a false premise, and the final answer is still wrong. Committing early carries a core risk that teams underestimate most as a failure mode, because everything looks fine until the last step lands somewhere useless.

Plan staleness is the related failure. The plan starts from valid assumptions, but conditions on the ground shift while execution is underway, and the Executor keeps faithfully carrying out a plan the world has since invalidated. Nothing in the base pattern notices this unless someone built a check for it.

The fix is a re-plan gate, and it has to be added deliberately: after every few steps, or whenever a step's output falls below some confidence threshold, the system asks the Planner whether to revise what's left. This falls outside the pattern by default. It's a decision someone has to make and implement. Context matters differently here, too. Whatever the Planner knew, or didn't know, at the moment it wrote the plan gets baked into every downstream step, so stale information at that single planning moment shapes the entire run.

Plan-and-Execute resists drift better than ReAct does, because the plan holds the goal steady across steps. But it resists surprise worse, since one unexpected tool result can invalidate everything downstream with no built-in correction. Adding re-planning fixes that, but it adds its own set of decisions: when to trigger it, how much of the plan to throw out, whether to restart execution from scratch. Each of those choices has a cost, and getting them wrong has a correctness price too.

How the two patterns compare on token cost, latency, and task accuracy

Diagram: ReAct vs. Plan-and-Execute: Cost, Latency, and Accuracy at a Glance. Visualizes: Show a side-by-side comparison of ReAct and Plan-and-Execute across three performance dimensions: latency (LLMCompiler-style Plan-and-Execute reports up to…

The cost shapes differ predictably. ReAct pays for roughly one call to a strong model per step, and the context it sends grows with every iteration. Plan-and-Execute pays for one strong-model call at planning time and then a series of cheap-model calls during execution. LangChain has noted that this split produces real cost savings over ReAct on sufficiently long tasks.

LLMCompiler (Kim et al., 2024, arXiv:2312.04511) pushes the planning approach further by scheduling independent steps to run at the same time instead of one after another. Reported maximum gains were 3.7x on latency, 6.7x on cost, and roughly 9% on accuracy compared with ReAct, on the workload the paper tested. Those are the best numbers observed for that setup, not numbers any team should expect by default.

Basic Plan-and-Execute without concurrent scheduling can actually run slower than ReAct, because it adds planning overhead without gaining anything back in parallelism. The latency win only shows up once the system identifies steps that don't depend on each other and runs them at the same time. AdaPlan, a Plan-and-Execute variant described in the PilotRL paper (arXiv:2508.00344), reports outperforming ReAct by 12.76% on complex agent tasks. A separate practitioner comparison published on dev.to put ReAct's task completion accuracy at 85% against Plan-and-Execute's 92% on the workload it tested. Treat that as a directional data point from one comparison, not a settled industry number.

Framework choice adds its own delta on top of pattern choice. Independent benchmarking cited in Intuz production data has LangGraph running roughly 2.2x faster than CrewAI on identical tasks, while LangChain's default pattern tends to re-send accumulated history at each step, which can increase token usage. Real infrastructure numbers from Intuz, drawn from over 100 deployments, put the monthly cost of handling 1,000 requests a day somewhere between $63 and $171. Where a given system lands in that range has more to do with which pattern it's running than which model.

How to read a task and decide which pattern fits

One question does most of the work: how much does the result of one step change what the next step should be?

If the answer is "a lot," if the right next move genuinely depends on live observations, ReAct fits, because its feedback loop is the whole point. If the dependencies between steps are stable and knowable ahead of time, Plan-and-Execute fits better, because there's nothing to gain from re-deciding the path at every turn. Teams that default to Plan-and-Execute for exploratory work are often paying planning overhead for a task that needed adaptability instead.

ReAct is the right default for exploratory or investigative work, where the right path only becomes visible after tools start returning results. It also suits short chains, where token accumulation hasn't yet become expensive, and real-time interactive settings, customer service, live query answering, where the latency of upfront planning isn't acceptable to a waiting user. Environments that require a per-step audit trail benefit from ReAct's granular step-by-step logging.

Plan-and-Execute earns its place on complex, multi-step tasks where the dependencies between steps are already known: research pipelines, structured data analysis, report generation. It's also the better fit anywhere a human or an automated reviewer should look at the plan before anything executes, which matters most for high-stakes or irreversible actions. Cost-sensitive systems running at real scale benefit from the planner-once, executor-cheap split, and tasks with genuinely independent steps can get a latency payoff through concurrent scheduling, as demonstrated by LLMCompiler-style variants.

When the choice still isn't obvious, the deciding question becomes which failure mode is easier to live with: wasted tokens from a ReAct loop that circles, or the added complexity of building re-planning logic for Plan-and-Execute. Neither answer is free. And neither pattern says anything about autonomy level on its own: both can run fully autonomous, and both can run as a copilot with a human checkpoint built in.

Adjacent patterns that extend or combine the two approaches

ReWOO (Xu et al., 2023, arXiv:2305.18323) takes the planning-first idea further than Plan-and-Execute does. The planner writes every step in a single pass, using variables as placeholders for results it hasn't seen yet. All the tool calls then run, potentially at the same time, and a synthesizer assembles the final answer at the end. The whole run costs only two LLM calls total, which the paper reports as roughly five times more token-efficient than ReAct on comparable tasks. The tradeoff is fragility: if any tool returns something the planner didn't anticipate, there's no mid-run reasoning step to catch it, because none exists.

Reflexion (Shinn et al., 2023) works differently. It adds a self-critique layer on top of whatever base pattern is running underneath. After an attempt, the agent critiques its own output and retries with that critique held in memory. On HumanEval coding tasks, this pushed pass rates from 80% to 91%, a real gain, though every retry means a full additional run, so the expense adds up fast. A 2025 replication study found a specific weakness in single-agent versions: because the same model both produces the output and writes the critique, it tends to repeat its own earlier misconceptions instead of correcting them, reinforcing rather than fixing its blind spots.

The native hybrid mentioned earlier, Plan-and-Execute at the strategic level with a ReAct loop inside each Executor step, gives each sub-task room to adapt while keeping the overall mission on a fixed track. It's the structure behind some of the more advanced deep research agents running today. Reflexion isn't tied to any one of these; it can sit on top of ReAct, Plan-and-Execute, or ReWOO as a correction layer rather than a base architecture of its own.

To restate it directly: ReAct, Plan-and-Execute, ReWOO, and Reflexion are patterns, ways of structuring reasoning. LangGraph, CrewAI, and AutoGen are frameworks that implement several of these patterns at once. LangGraph leans on stateful graphs to support re-planning, CrewAI handles tool access through declarative scoping, and AutoGen ships built-in sandboxed execution.

Security and production-hardening considerations that differ by pattern

Security is asymmetric across the two patterns, and the difference is structural. Del Rosario et al. (2025, arXiv:2509.08646, University of Oxford) found that Plan-then-Execute carries an inherent advantage here: by establishing control-flow integrity upfront, before any tool runs, it resists indirect prompt injection better than ReAct does. In ReAct's loop, a malicious tool output landing in an Observation can redirect the very next Thought, and no plan sits above it to catch that redirection. Anyone running ReAct on a task that touches untrusted external content should treat that as a real exposure, not a theoretical one, and design around it rather than hope it doesn't come up.

The Verifier component, sometimes called Plan-and-Verify-Execute, is one practical answer. A separate model or a rule-based engine checks the plan before the Executor touches it, ideally using a different prompt or persona than the Planner used, so it isn't prone to repeating the same blind spots. That extra check earns its cost most clearly when a wrong execution would be expensive or impossible to undo.

Some hardening steps apply regardless of which pattern runs underneath. The principle of least privilege says the Executor should only get access to the tools its current step actually needs, nothing more. CrewAI's declarative tool scoping is a working example of enforcing that directly. Sandboxed code execution matters just as much: AutoGen's built-in Docker sandboxing addresses that concern head-on, containing whatever code an agent decides to run.

ReAct's adaptability cuts against it here too. Every new Observation is a fresh entry point for injected content, since the agent trusts and acts on whatever it reads, step by step, with no upfront plan standing between a bad input and the agent's next move.

Sources

  1. ReAct vs. Plan-and-Execute: Agent Architecture Guide [2026]
  2. Single-Agent Patterns
  3. Architecting Resilient LLM Agents: A Guide to Secure Plan-then-Execute Implementations
  4. ReAct vs Plan-and-Execute: A Practical Comparison of LLM Agent Patterns
  5. PilotRL: Training Language Model Agents via Global Planning-Guided Progressive Reinforcement Learning
  6. pub.towardsai.net

More in AI Agent Architectures