ProcessFrontier R&D
Is it the memory or the model? A protocol for agent memory evaluation
When an agent forgets what it agreed last week, teams tend to blame the model or the memory layer on instinct. This protocol replaces instinct with an experiment: a long-horizon task suite with known answers, recorded tool responses so runs can be replayed, ablations that change one component at a time, repeated runs to measure variance, and a taxonomy that sorts every failure by cause. The result is evidence about where to spend engineering effort.
On this page
- Model fault or memory-design fault: the attribution problem
- The evaluation protocol at a glance
- Running the memory study step by step
- An ablation matrix for memory, retrieval and model
- A failure taxonomy for agent memory and coordination
- Threats to validity in agent memory studies
- From finding to fix: the transfer package
- Questions and answers
- Sources
Model fault or memory-design fault: the attribution problem
An agent system that performs well in trials but loses context across sessions is one of the example questions behind our Frontier R&D practice1. The symptom, a forgotten commitment or a stale fact, can come from what was written to memory, how it was retrieved, how it was summarised or how the model used what it received. Each cause implies different work, and guessing wrong spends a quarter on the wrong fix.
Production traces cannot settle this, because the model, the data and the users all change between sessions and agents are not deterministic. A controlled study holds everything fixed except the component under test. Which architecture to choose is covered on our agent swarms page; this protocol is about measuring the one you have.
The evaluation protocol at a glance
- Task suite with ground truth
Multi-session scenarios whose correct recalled facts are known in advance.
- Record tool responses
Capture each external call once so later runs replay identical inputs.
- Fix and log configurations
Pin model versions, prompts, memory settings and seeds for every run.
- Ablate one component
Swap the memory store, retrieval policy, summariser or model in isolation.
- Repeat runs
Run each configuration several times to see the spread of outcomes.
- Classify failures
Tag every missed recall point with a cause from the taxonomy.
- Report with intervals
State effects with confidence intervals and threats to validity.
Running the memory study step by step
Build a long-horizon task suite
Write multi-session scenarios modelled on the production failure: facts introduced early, changed midway and needed later. Each carries the ground truth for every recall point, so scoring is mechanical. Include distractors such as similar customers or superseded instructions, because memory fails on near-misses.
Make every run reproducible
Record tool responses once and replay them, so a change in an external API cannot pose as a memory effect. Pin model versions, prompts, sampling settings, memory parameters and seeds, and log them with each run.
Reproduce the failure first
Run the current production configuration on the suite. If it does not fail the way production does, the suite is missing something and every later comparison is suspect.
Ablate one component at a time
From the baseline, change exactly one element per configuration. Test combinations later, and only for components that showed an effect on their own.
Repeat runs and estimate variance
Agents give different answers to identical inputs. Run each configuration enough times to see the spread, and report differences with confidence intervals rather than single scores.
Classify each failure and report
Tag every miss using the taxonomy below, then write up effects, intervals, threats to validity and the recommended change.
An ablation matrix for memory, retrieval and model
Each row changes one thing relative to the baseline. Reading the last column shows what a difference from the baseline would point to.
| Configuration | What changes | Held fixed | What a difference suggests |
|---|---|---|---|
| Baseline | Nothing; current production design | Everything | The reference profile for every other row |
| Memory store swap | Free-text summaries replaced by structured fact records | Model, retrieval, prompts | Lost or contradictory facts trace to how memory is written |
| Retrieval policy swap | Similarity search replaced by recency-weighted or filtered retrieval | Model, store, prompts | Facts exist in memory but are not surfaced when needed |
| Summarisation off | Raw transcripts stored instead of summaries | Model, store, retrieval | Detail is lost during compression |
| Model swap | Different model family or size | Store, retrieval, prompts | The model mishandles context it was given |
| Oracle memory | Correct facts injected directly into context | Model and prompts | An upper bound: if this still fails, memory is not the main problem |
The oracle row is often the most informative single experiment, because it separates a fact that never arrived from a fact the model could not use.
A failure taxonomy for agent memory and coordination
- Lost fact
- Information given in an earlier session is absent from memory when it is needed.
- Stale fact
- A superseded value is recalled, such as an old delivery address.
- Contradictory memory
- Memory holds conflicting versions of a fact and the agent picks the wrong one or hedges.
- Retrieval miss
- The correct fact is stored but retrieval does not return it for the current step.
- Use failure
- The correct fact reaches the model's context but the response ignores or misapplies it.
- Coordination failure
- In a multi-agent system, one agent holds the fact but never passes it to the agent that needs it, or two agents write conflicting updates.
Threats to validity in agent memory studies
The suite does not resemble production
Early signalThe baseline passes scenarios that fail in production.
MitigationRebuild scenarios from de-identified failure traces before trusting any comparison.
Too few repetitions
Early signalRankings between configurations flip when the study is rerun.
MitigationAdd runs until intervals clearly separate or clearly overlap, and report either outcome.
Hosted model drift
Early signalProvider-side behaviour changes during the study window.
MitigationPin versions where possible and rerun the baseline at the end to detect drift.
Scoring errors
Early signalAutomated scores disagree with human spot checks.
MitigationScore against answer keys where possible and audit a sample of model-graded items by hand.
From finding to fix: the transfer package
The finding is rarely just "change the retriever". A useful package contains the suite and replay fixtures, ablation results with intervals, failure counts by cause, the recommended change with its evidence, and the configurations tried and rejected, because someone will propose those again.
The suite then becomes a regression test: rerunning it after each model upgrade shows whether the fix still holds. Release gates are a separate discipline, covered in our agent production readiness process; this study answers the earlier question of what to build.
Questions and answers
How many scenarios does a useful agent task suite need?
Enough to exercise each failure type several times, with distractors, rather than a fixed count. A small suite built from real failure traces usually teaches more than a large synthetic one. Start with scenarios that reproduce the production problem, confirm the baseline fails on them, then add variations until failure counts by cause stay stable between reruns.
Can we use a public agent benchmark instead of building our own suite?
Public benchmarks help compare models in general, but they rarely contain your domain's facts, distractors or session patterns. Use one as a sanity check alongside a suite built from your own traces. If the two disagree, trust the one that resembles production, and treat the disagreement as a finding worth explaining.
Why record tool responses instead of calling live systems?
Live systems change between runs: prices move, records update and APIs time out. Any of those can alter an agent's behaviour and be mistaken for a memory effect. Replaying recorded responses gives every configuration an identical world, which is the only way a difference can be attributed to the component that changed.
Does this protocol work for multi-agent systems?
Yes, with additions. Each agent's memory and the messages between agents become components that can be ablated, and the coordination category of the taxonomy matters more. Log which agent held each fact and when it was handed on, so a missed recall can be traced to the agent that lost it or the hand-off that never happened.