Agent Memory Types in Long-Running Tasks
Persistent memory layers, not bigger context windows, solve agent degradation.

Long-running AI agents don't degrade because the underlying model gets worse at reasoning. They degrade because nothing about the system carries information forward from one moment to the next, and the industry keeps trying to fix that with bigger context windows instead of better memory design.
In-context memory: what the agent is actually reasoning over right now
In-context memory is everything the model actually looks at during one pass: the system prompt, the conversation so far, whatever got pulled in from a search, and any tool output sitting in the buffer. It's the only memory type the model directly reasons over. Every other kind of memory, whatever happened last week, whatever fact got learned last month, has to get pulled into this window before it can shape a single word of output.
The comparison to human short-term memory holds up better than most AI analogies. Alan Baddeley's model of working memory splits things into a central executive and a set of buffers, and that maps onto an LLM agent almost exactly: the model is the executive, the context window is the buffer, and both run into the same wall. There's only so much you can hold in mind at once, no matter how sharp the mind is.
Context windows have gotten big. GPT-5.2 handles somewhere around 400,000 tokens, and Qwen2.5's long-context variant runs to 100,000. Those numbers sound like a fix for the persistence problem, and they help, but they don't solve it. A big window still resets. Close the session, open a new one, and none of that accumulated context survives unless something outside the model saved it. Size buys you more room in the moment, not memory across moments.
Even within a single long window, the model doesn't treat every token equally. Research on long-context retrieval, including work behind the LongBench benchmark, found that models reliably use information near the start and end of a long context but underuse whatever sits in the middle, a pattern usually called "lost in the middle." So stuffing more tokens into the prompt doesn't guarantee the model reasons well over all of them. It just gives you more places for something important to get buried.
That leaves a design question, not a scaling question: what goes into context, when it goes in, and where it gets pulled from. Answering that well is the actual engineering problem. Bigger windows just raise the ceiling on how badly you can get it wrong.
Episodic memory: preserving what happened so the agent can learn from it
Episodic memory holds records of specific things that happened: a tool call, a turn in a conversation, an observation from the environment, an outcome, usually tagged with a timestamp and some measure of how important it was. This is the layer that lets an agent say "the last time this came up, here's what happened" instead of treating every interaction as the first one.
The clearest demonstration of what this buys you came from the Generative Agents work by Park and colleagues in 2023, where 25 simulated agents ran for 48 hours in a small virtual town. Each agent kept a three-tier memory: raw observations, higher-level reflections distilled from those observations, and a retrieval step that pulled back whatever was both recent and important. Out of that setup, agents threw parties, formed relationships, and passed along information, behavior nobody scripted directly. It came from accumulated episodes, not from a bigger model underneath.
MemGPT, from Packer and colleagues, framed the same problem in systems terms: treat the model's active context like RAM and an external store like disk, and let the agent page information between the two as needed. It's a clean analogy because it names the actual bottleneck. The model can only hold so much live at once, so something has to manage what moves in and what gets archived.
Where current systems still struggle is time itself. The Episodic Memory Benchmark, built specifically to test whether agents track how things change over a sequence of events, scored the best model tested, Gemini-2-Pro, at 0.290 on its Chronological Awareness measure, under 30% accuracy at correctly sequencing events. Agents are decent at recalling that something happened. They're much worse at recalling when it happened relative to everything else.
Most production systems handle this with vector databases and semantic search, retrieving episodes that are conceptually close to the current situation rather than requiring an exact match. That's useful and it's also the limit. An episodic system, by itself, can't turn ten instances of the same correction into one general rule. It just keeps stacking up instances. Getting from a pile of episodes to a usable rule is a different job, handled by a different layer.
Semantic memory: turning repeated experience into reusable knowledge
Semantic memory strips the context away and keeps the fact. Not "the user corrected the date format on January 5th," but "the user prefers DD/MM/YYYY." Not a record of an event, but a conclusion drawn from several of them, filed as knowledge the agent can apply the next time the topic comes up regardless of which past episode triggered it.
Getting from episodic to semantic is supposed to be a consolidation step, and in most systems deployed today, that step doesn't happen on its own. Somebody has to write a prompt that says "summarize what you've learned about this user" or set up a heuristic trigger that fires the consolidation after some number of repeated instances. Autonomous consolidation of this kind, where an agent notices a pattern across episodes and quietly promotes it to a standing fact, remains an unsolved challenge in practice.
It helps to think of semantic and procedural memory not as separate boxes sitting next to episodic memory, but as different stages of processing the same raw material: experience turns into knowledge, and knowledge, eventually, turns into skill.
On the storage side, teams tend to reach for one of two patterns. Knowledge graphs handle entity-relationship reasoning well, things like who reports to whom, or which product depends on which service. Vector databases paired with retrieval-augmented generation (RAG) pipelines handle looser, unstructured domain knowledge. RAG's weak point is baked into how it works: it retrieves based on how similar a chunk's embedding is to the query's embedding, and similarity in that mathematical sense doesn't always line up with what's actually useful for solving the problem in front of the agent. The two things, embedding similarity and task usefulness, are optimized separately, not together, and that gap shows up as retrieved chunks that are on-topic but not helpful.
There's also a subtler failure worth naming: semantic drift. An agent forms a nuanced understanding of something, that understanding gets chunked and embedded by the memory system in a way that doesn't preserve the nuance, and the next time the agent retrieves it, what comes back is a fragment that's only loosely related to what was actually meant. The agent's internal sense of what it knows and what the memory system actually stored start to pull apart.
One pattern that's gained traction as a practical fix is scoping memory writes at the point of storage: tagging each entry with a user ID, an agent ID, a session ID, an organization ID, and then composing and ranking across those scopes at retrieval time. That lets a company-wide policy fact and one user's personal preference sit in the same memory store without stepping on each other.
Procedural memory: encoding how to do things, not just what is known
Procedural memory is the odd one out among the four, because it's not about facts at all. It's about method: a verified sequence of actions that worked before and can be run again. Voyager, an agent designed to play a game, is the cleanest working example. Every skill it verified got saved as executable code, indexed by a plain-language description, and pulled back out and recombined whenever a new task called for it.
That's procedural memory working the way it's supposed to: as code, executable and unambiguous. Most procedural memory in deployed agents today isn't stored that way. It's stored as text, a written-out description of a workflow, and that's where things go sideways. An agent can retrieve a perfectly accurate description of a ten-step process and still fail to execute it correctly, skipping a step, doing two steps out of order, because knowing the procedure in words and instantiating it as a sequence of actions turn out to be two different skills. Declarative facts ground reasoning just fine in plain text. Procedures described in plain text have a harder time surviving the trip from description to execution.
There's active research trying to close that gap by moving procedural memory below the level of tokens entirely, encoding it directly into the model's internal activations rather than as a retrievable block of text, an approach sometimes described under the banner of neural procedural memory. It's early, but it points at something real: recalling a fact reliably is a largely solved problem at this point, while representing and using a procedure reliably is not. That gap is one of the more stubborn open problems in the field.
More recent systems have started tackling the coordination problem directly rather than the representation problem alone. AdMem, a framework out of work associated with Princeton, Amazon, and Arm, wires semantic, episodic, and procedural memory into one structure split between short-term and long-term stores, run by a small set of cooperating agents (an actor, a memory agent, and a critic) that merge and prune memories based on which ones actually paid off. It's a direct answer to a problem earlier procedural-memory systems didn't handle well: keeping the whole thing usable as it grows, rather than accumulating indefinitely until retrieval slows to a crawl.
How the four types form a reasoning stack, and where the orchestration breaks
Picture an agent processing a customer return. Procedural memory supplies the steps: check eligibility, verify the receipt, issue the refund. Semantic memory supplies the policy itself: the return window, which categories of item qualify. Episodic memory supplies the specific history: what this customer did last time, whether there was a dispute. In-context memory is where all three get held together long enough for the model to actually reason across them and produce a decision.
That's the stack in theory. In practice, the hardest unsolved question in the whole field is the transition policy between layers. When does a specific episode get promoted into a general semantic fact? When does a stored semantic fact get pulled back down into working memory for a task that needs it right now? Almost every deployed system answers these questions with a hardcoded rule or an explicit prompt telling the agent to check. Nobody has a reliable, autonomous mechanism that consolidates memory on its own the way it needs to.
Failure shows up in both directions. Page too much into context and you burn tokens on history that has nothing to do with the current task, crowding out the parts that matter. Archive too aggressively and you get what's sometimes called memory blindness: the fact the agent needs is sitting in cold storage, technically retrievable, but the agent has no signal telling it the fact exists at all, so it never asks.
Multi-agent setups add a version of this problem that's easy to miss. One agent stores a finding along with the reasoning behind it. A second agent retrieves that finding later, but gets only the conclusion, not the reasoning that produced it. Both agents technically share the same data. They don't share the same understanding of what that data means, and that gap tends to surface at exactly the moment it matters most.
Recent survey work on agent memory names all of this directly: consolidation from episodic to semantic, the transition of memory into something closer to the model's own implicit knowledge, and governance of memory across multiple agents are all flagged as open, unresolved research problems, not solved engineering ones. That body of work is less a fixed architecture than a description of the job that still needs doing.
How current memory systems are actually built and where each approach hits its limit
Three broad approaches dominate right now, and each one solves part of the problem while leaving another part exposed.
Long-context models try to solve memory by making the window big enough that persistence looks automatic. GPT-5.2's roughly 400,000-token window is the current high end of this approach. It runs into the lost-in-the-middle effect described earlier, and more fundamentally, it doesn't survive past the session. Nothing in a long context window carries over once that window closes.
RAG-based external memory takes the opposite approach: push information out of the model entirely, into a vector store, and pull relevant pieces back in at query time based on embedding similarity. It's the most widely deployed pattern in production today, and it works reasonably well for a lot of use cases. Its limit is structural: relevance by embedding distance and usefulness for the actual task aren't the same measure, and the pipeline isn't trained end-to-end to close that gap.
Agent-centric memory management is the newer approach, where the agent itself decides what to store, how to structure it, and when to retrieve it, building up its own distilled priors from its own task history rather than relying entirely on an external retrieval system to guess what's relevant. This is the approach best suited to tasks that run long, span many steps, and touch real-world systems where conditions keep changing.
Worth separating here: agent-centric memory (what the agent has learned about doing its job) and user-centric memory (what the agent has learned about the specific person it's working with) aren't competing designs. Production systems generally need both running at once.
There's also a split between where memory lives. Client-side memory is private to one agent and managed directly by it. Server-side memory sits with a provider and supports things a single agent can't do alone: shared learning across multiple agents, enterprise-grade data governance, audit trails, version history, and attribution for where a piece of knowledge came from. Red Hat's 2026 architecture writeup on this topic makes the point plainly: memory and harness design can matter as much as raw model capability, which says more about the value of the surrounding system than about the model itself.
Realistic long-running tasks, the kind involving evolving state, feedback logs, half-finished plans, and a trail of prior actions, need memory that's structured and queryable, not just present. Skip that and agents repeat mistakes they already made, forget approaches that already worked, and give inconsistent answers to questions they've already answered correctly once.
Gartner's 2026 outlook makes a related, sharper claim: agentic analytics projects that rely only on a model connectivity standard, without a proper semantic layer underneath, face a 60% failure rate by 2028. That's a direct statement that no single layer, however well built, covers the whole problem.
What current benchmarks reveal about where agent memory actually works and where it does not
Three benchmarks currently anchor how the field measures progress. LoCoMo runs 1,540 questions spanning single-hop lookups, multi-hop reasoning, open-domain questions, and temporal recall. LongMemEval runs 500 questions covering categories like tracking updated knowledge and pulling information across multiple sessions. BEAM pushes further out, testing at context scales of one million and ten million tokens across several task categories. These approaches track token cost alongside accuracy, treating efficiency as a real scoring dimension rather than a footnote.
Mem0's 2026 results give a concrete sense of what "good" currently looks like: 92.5 on LoCoMo and 94.4 on LongMemEval, using under 7,000 tokens per query on average (6,956 on LoCoMo, 6,787 on LongMemEval). The gains over prior systems were sharpest exactly where earlier benchmarks flagged the weakest spot: +29.6 points on temporal reasoning and +23.1 points on multi-hop reasoning. That temporal-reasoning gain lands on the same dimension where Gemini-2-Pro scored 0.290 on the Episodic Memory Benchmark's chronological test, which suggests the gap is closable with the right architecture rather than being some inherent limit of current models.
The 6,900-tokens-per-query figure matters on its own. It's a rough price tag for what near-95% accuracy actually costs once memory architecture is done well, and it gives anyone evaluating a system something concrete to check a vendor's numbers against.
None of this means the measurement problem is solved. The the same body of survey work that flagged consolidation and governance as open questions also treats benchmarking itself as underdeveloped: current tests don't do a good job capturing dynamic environments, coordination across multiple agents, or an agent's ability to keep learning over an extended, continuous run rather than a fixed test set. Newer work is pushing directly at that gap: some recent systems apply reinforcement learning at runtime over episodic memory to let agents improve their own memory use as they go; "Memory as Action" (November 2025) treats the decision of what to keep in context as something the agent itself should actively manage rather than a fixed retrieval step; other recent work ties planning and retrieval together rather than treating them as separate stages. None of these are settled yet. They're the direction the field is currently pushing.
What this architecture means for teams building or buying agents at scale
Gartner projects that by the end of 2026, 40% of enterprise applications will have task-specific AI agents built into them, up from under 5% in 2025. That's a fast climb, and it means the memory architecture decisions being made right now, often by teams under deadline pressure, will end up governing how most of these deployments actually behave once they're running in production rather than in a demo.
The governance angle isn't separate from the engineering angle. Gartner's projection that half of enterprise agent deployment failures by 2030 will trace back to weak runtime enforcement of governance controls is really a statement about memory design. A memory system with no audit trail, no version history, and no scope controls between users and organizations isn't a rough edge to smooth over later. It's a failure mode waiting for the right conditions to trigger it.
Anyone evaluating an agent system, whether building it or buying it, has a short list of questions worth asking before anything else. Is working memory getting loaded with what actually matters, or is it filling up with history that crowds out the useful signal. Is there any path from repeated episodes to a standing rule, or does every session start from zero. Does procedural knowledge exist as something the agent can execute, or only as a description the agent has to reinterpret every time. And is there any accounting for who wrote which memory, when, and under what authority, for the moment someone needs to explain why the agent did what it did.
The model matters less here than the four systems built around it. That's not a hedge. It's the whole argument.
Sources
- Beyond Short-term Memory: The 3 Types of Long-term Memory AI Agents Need - MachineLearningMastery.com
- From context to dreams: architecting memory for AI agents
- A Survey of Agent Memory in the Second Half: Towards Self-Evolving and Long-Horizon Agents
- AdMem: Advanced Memory for Task-solving Agents
- State of AI Agent Memory 2026: Benchmarks & Trends Report
- Memory as Action: Autonomous Context Curation for Long-Horizon Agentic Tasks


