The Context Window

Prompt Caching Economics in High-Volume Agent Systems

Clever prefix design and TTL timing determine whether agents actually capture caching savings.

Staff Writer · · 11 min read
Cover illustration for “Prompt Caching Economics in High-Volume Agent Systems”
Context Delivery · October 2, 2026 · 11 min read · 2,513 words

Prompt caching now functions as a major cost lever available to teams running large language models in production, and the discount it advertises is aggressive: cache reads cost roughly a tenth of base input at both Anthropic and OpenAI, and between 75–90% less at Google Gemini 2.5 Flash. This piece lays out what the mechanism does, where the savings quietly disappear in agentic systems, and the specific design choices, prefix layout, tool routing, and keepalive timing, that decide whether a high-volume deployment pays full price or a fraction of it.

How prompt caching works and what it costs across providers

Caching works by storing the key-value state a model computes while reading a prompt prefix; when a later request starts with that same prefix, the provider serves the stored state instead of recomputing it, and the model skips the prefill step for those tokens entirely. Anthropic, OpenAI, and Google have each built this into standard pricing, and the discount on a cache hit is substantial everywhere it's offered: roughly a tenth of the base input price at both Anthropic and OpenAI, and a 75 to 90 percent reduction at Google on Gemini 2.5 Flash.

The three providers diverge sharply on what it costs to write to the cache in the first place, and that divergence shapes which workloads actually benefit. Anthropic charges a real premium to write into its shorter cache tier and doubles that premium for the 1-hour tier, while GPT-5.5 carries no write premium at all, and Gemini's implicit cache charges nothing extra to write either. How long a cached entry survives before eviction varies just as widely: Anthropic's default cache tier is short-lived, OpenAI's older in-memory cache clears after a brief idle window, GPT-5.5 and later default to 24-hour extended retention, GPT-5.6 and later use an explicit minimum TTL tied to defined breakpoints, and Google's implicit cache does not reliably evict at all.

That landscape shifted meaningfully on July 9, 2026, when OpenAI rolled out GPT-5.6 and replaced its free, implicit caching model with explicit breakpoints, a 1.25× write premium, and a 30-minute minimum TTL, pulling OpenAI's pricing structure much closer to Anthropic's and ending the period in which caching on that platform cost nothing to write. None of these provider differences make one platform categorically superior to another; they mean the same architectural choice produces different economics depending on which provider is in use. Every argument that follows in this piece rests on that premise: the provider sets the rules, but the system design decides how much of the discount actually gets collected.

Diagram: Cache Read Discounts vs. Write Premiums Across Providers. Visualizes: Show the asymmetry between cache read discounts and write premiums across three providers — Anthropic, OpenAI (GPT-5.6), and Google Gemini 2.5 Flash.

Why the 90% discount is not automatic

A cache read discount only applies when the prefix a request sends matches a recently cached prefix exactly, and that condition fails far more often in practice than most engineering teams expect going in. A KV-cache-aware prompt engineering benchmark from August 2025 measured this directly: requests with stable, unchanged prefixes cost far less per call than requests whose prefixes had been perturbed, even when the underlying query content was identical, so the cost difference tracked prefix stability alone, not what the user actually asked. That benchmark result is specific to the study that produced it, not a general claim about the industry, but it demonstrates the mechanism that every later section in this piece depends on: a cache either sees the same bytes it saw before, or it doesn't, and there's no partial credit.

Anthropic's cache has a two-tier architecture with a sharp threshold in the low thousands of tokens, and research (Song, PayPal, arXiv 2607.15516, July 2026) found that below this threshold the hit rate plateaus well below the ideal, meaning short prefixes do not behave as the pricing model implies. A workload built on prompts that look stable on paper but produce a hit rate far below expectations is telling its engineers something concrete: some field inside the supposed-static prefix is still changing between calls.

The arithmetic of when caching pays off is specific to each provider's tier structure. On Anthropic's 5-minute cache tier, the system breaks even at roughly 1.4 reads per cached write, so a workload with a stable prompt but a low reuse rate can end up paying more in write premiums than it recovers in read discounts. The 1-hour tier, which carries a 2× write premium, breaks even at fewer than 3 reuses per hour, a bar most production agent workloads clear without difficulty, provided the prefix genuinely holds steady for the full hour. That proviso is where most of the loss actually happens, because a system that pings a cache entry after it has already expired doesn't refresh anything: it re-prefills a dead entry at full price, paying the write premium a second time for no benefit. Whether the discount materializes depends on whether the engineering around that call respects the provider's specific tier math and TTL window.

How agentic loop structure systematically destroys cache residency

Autonomous agents follow a think-act-wait pattern that is structurally hostile to caching: the system sends a request, waits on a tool, a build, a test suite, a deployment, or a human approval, and then sends a follow-up that would reuse the large conversation prefix if that prefix were still in the cache. The trouble is timing. Provider TTLs and the pause an agent takes waiting on a tool call are both measured in minutes, and the pause is the call that would reuse the large conversation prefix, but the cache frequently loses. At the scale a real agent task runs, dozens of pauses across a single task, prefixes stretching into the tens of thousands of tokens, this is a recurring cost line that appears on every invoice.

Tool-heavy agents make the problem worse before they make it better, because system prompts and tool definitions can take up a large share of the token budget on every request, and caching that block only pays off if it stays byte-for-byte stable across the entire loop. ProjectDiscovery's Neo platform supplies the clearest documented case of this failure in production. Neo is an autonomous security testing system running multi-agent, multi-step workflows that routinely execute 20 to 40 or more LLM steps per task, with system prompts running to thousands of lines of YAML and tens of thousands of tokens per agent, and every step re-sending the full conversation. According to ProjectDiscovery's own engineering blog, Neo ran in production at a 7 percent cache hit rate despite caching being reported as enabled, and the cause traced back to mutable working memory sitting inside what was supposed to be a static prefix, so every single step invalidated the entire cached block that sat behind it.

The lesson generalizes well past Neo's specific codebase. Any field that changes from one step to the next, a timestamp, a counter, session state, a per-call identifier, placed anywhere before the cache breakpoint will silently invalidate that cache on every request, and the system will keep paying full price while appearing to have caching turned on. This is the failure mode every architectural fix described in the rest of this piece exists to prevent.

Prefix layout and breakpoint placement as the primary cache design decision

Diagram: Static Before Dynamic: The One Rule That Decides Cache Residency. Visualizes: Visualize the correct prompt prefix ordering that determines whether a cache hit occurs.

The single most consequential decision in cache architecture is where the breakpoint sits in the prompt, because everything before that line has to be genuinely static across every call, and everything that changes from request to request has to live after it. The correct ordering runs from the static system prompt, through static tool definitions, through static few-shot examples, through cached document context, and only then into dynamic per-request content such as the user's query, the current conversation turn, or working memory. Placing any mutable element ahead of that line invalidates the entire cached block on every single call, exactly as Neo's production incident demonstrated.

How a team sets that breakpoint differs by provider in ways that carry real engineering weight. Anthropic exposes explicit cache_control markers, so a developer chooses precisely where the breakpoint falls and pays the corresponding write premium for each cached block. OpenAI's pre-GPT-5.6 models worked differently: caching applied implicitly above a certain token threshold, which meant no explicit marker was required, but it also meant no control over what got cached. GPT-5.6's shift to explicit breakpoints closes that gap, and OpenAI deployments now need the same discipline around prefix placement that Anthropic teams have always had to practice.

Breakpoints aren't free to add. Each additional cache tier a request uses carries its own write premium, so the real design question is how many tiers a given workload can justify, and a single breakpoint covering the whole static prefix is usually more efficient than several nested breakpoints unless each inner tier earns enough reuse on its own to be worth its premium. The CAPC research introduces a tier-preserving ratio bound that prevents over-compression from pushing the cached prefix into the lower-hit-rate tier, a concrete demonstration of how breakpoint placement interacts with cache tier architecture.

This architecture inverts a habit most engineers learned before caching existed. The old incentive was to keep system prompts short, since every token in the prompt added directly to input cost. Caching reverses that logic: a very large system prompt that holds a high cache hit rate across repeated calls costs less in practice than a short prompt that never gets cached at all, because the long prompt pays its write premium once and then bills at the cache-read rate on every subsequent call, while the short uncached prompt pays full input price every time. Teams optimizing prompt length for cost reasons, rather than optimizing for hit rate, are solving the wrong problem.

The tool definition problem and the dual-path architecture that resolves it

Tool definitions sit near the front of most agent requests, which puts them squarely inside the static prefix a cache depends on, and that creates a direct conflict with a popular design pattern: progressive tool disclosure, where an agent shows only the tools relevant to its current step in order to keep context small. Every time the visible tool set changes between calls, the prefix changes with it, and a cache that requires a fixed prefix cannot tolerate that kind of churn.

The obvious workaround, sending every tool definition on every call regardless of relevance, keeps the prefix stable, but it comes at a real cost: a 2026 analysis found that system prompts and tool definitions together can consume 40 to 60 percent of a tool-heavy agent's token budget, and that share grows in direct proportion to how many tools the agent supports. A system built this way stabilizes its cache hit rate at the price of bloating every single request with definitions the current step will never touch.

CacheRouter, proposed by Zha and colleagues (arXiv 2608.22708), resolves the conflict by splitting tool access into two separate paths rather than forcing one channel to do both jobs. The main model only ever sees a small, fixed core set of tools, which keeps its prefix stable, while every other tool runs through an independent routing channel: a router sub-model selects the right tool, executes it, and returns only the result to the main model's context. Tool registration pulls automatically from source code and supports runtime updates, so the available tool set can keep growing without ever touching the main model's request prefix. In a prototype tested against dozens of functional queries and a multi-turn dialogue, this design achieved very high token-level cache hit rates and cut input cost to a small fraction of a no-cache baseline under DeepSeek's pricing, though that result describes the prototype's measured performance rather than a claim about production deployment at scale.

The architectural principle behind CacheRouter generalizes past the specific paper: tool selection, deciding which capability a step needs, has to be kept separate from tool delivery, deciding which definition actually enters the model's request. Collapsing those two functions into a single channel means every change in selection rewrites the prefix and forces a fresh cache write. The Model Context Protocol serves as CacheRouter's standardized tool description layer, and the dual-path pattern it demonstrates extends to any runtime built on that same protocol.

Keepalive strategy for preserving cache residency across agentic pauses

Fixing prefix layout solves one part of the problem that agentic loops create; keeping a correctly structured cache entry alive through the pause between a request and its follow-up solves the other. A client-side keepalive, a lightweight ping that replays the cached prefix while a tool runs, is an individually sound tactic and a measurably effective one, but research shows the frequency most teams default to is far higher than it needs to be.

The paper's most useful finding runs counter to the 30-second convention: keepalive cost falls steadily as the interval between pings grows, so the economical choice is the longest interval that still stays safely under the provider's TTL.

TTLs vary sharply by provider: Anthropic's default cache is short-lived, OpenAI's legacy in-memory cache clears after a brief idle period; GPT-5.5+ defaults to 24-hour extended retention; GPT-5.6+ uses an explicit minimum TTL with defined breakpoints, and Google's implicit cache does not reliably evict. On OpenAI, now running an explicit 30-minute TTL since GPT-5.6, keepalive pings deliver real savings at a 30-minute pause. On DeepSeek, where the TTL runs around 10 minutes, re-prefilling a cold entry is cheap enough that a keepalive buys lower latency but not meaningful cost savings. On Google, the cache essentially never evicts on its own, so a keepalive there offers a latency benefit without a cost benefit attached to it.

The mechanism relies on reading a cached prefix resetting its position in the provider's eviction ranking, so a periodic read is what keeps the entry alive rather than any special "keepalive" API. Spacing those reads correctly, near the top of the safe interval rather than far below it, is what turns a defensive tactic into an actual cost saving rather than an unnecessary expense layered on top of an already-working cache.

When prompt compression and prompt caching

Prompt compression and prompt caching solve adjacent but distinct problems. Compression reduces the number of tokens a request sends; caching reduces the cost of tokens a request sends repeatedly. In that scenario, the compression succeeds on its own terms, fewer tokens per call, while quietly destroying the far larger saving the cache was already providing.

The two techniques work together cleanly only when applied to different parts of a request: compression on the dynamic content that changes every call, where there's no cache benefit to protect in the first place, and caching on the static prefix, left at whatever length it needs to be to do its job. Shrinking a system prompt that already holds a high cache hit rate rarely saves meaningful money, since the bulk of its tokens are already billing at the discounted read rate; shrinking the dynamic query or working-memory block that sits after the breakpoint saves real tokens without touching anything the cache depends on. Know which block a technique is acting on before applying it, because the static prefix and the dynamic tail follow different economics.

Sources

  1. Keeping the Cache Warm Pays: Keepalive Economics for Agentic Workloads
  2. CacheRouter: A Dual-Path Tool Routing Architecture with Cache-Preserving Main-Model Isolation for Long-Tail Tool Discovery
  3. Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching
  4. Prompt Caching Economics: Designing Cache-First Agents
  5. Don’t Break the Cache: An Evaluation of Prompt Caching
Filed underContext Delivery

More in Context Delivery