ProcessFrontier R&D

Is it the memory or the model? A protocol for agent memory evaluation

When an agent forgets what it agreed last week, teams tend to blame the model or the memory layer on instinct. This protocol replaces instinct with an experiment: a long-horizon task suite with known answers, recorded tool responses so runs can be replayed, ablations that change one component at a time, repeated runs to measure variance, and a taxonomy that sorts every failure by cause. The result is evidence about where to spend engineering effort.

Reviewed 6 min read

On this page
  1. Model fault or memory-design fault: the attribution problem
  2. The evaluation protocol at a glance
  3. Running the memory study step by step
  4. An ablation matrix for memory, retrieval and model
  5. A failure taxonomy for agent memory and coordination
  6. Threats to validity in agent memory studies
  7. From finding to fix: the transfer package
  8. Questions and answers
  9. Sources

Model fault or memory-design fault: the attribution problem

An agent system that performs well in trials but loses context across sessions is one of the example questions behind our Frontier R&D practice1. The symptom, a forgotten commitment or a stale fact, can come from what was written to memory, how it was retrieved, how it was summarised or how the model used what it received. Each cause implies different work, and guessing wrong spends a quarter on the wrong fix.

Production traces cannot settle this, because the model, the data and the users all change between sessions and agents are not deterministic. A controlled study holds everything fixed except the component under test. Which architecture to choose is covered on our agent swarms page; this protocol is about measuring the one you have.

The evaluation protocol at a glance

01Task suite with groundtruth02Record tool responses03Fix and logconfigurations04Ablate one component05Repeat runs06Classify failures07Report with intervals
  1. Task suite with ground truth

    Multi-session scenarios whose correct recalled facts are known in advance.

  2. Record tool responses

    Capture each external call once so later runs replay identical inputs.

  3. Fix and log configurations

    Pin model versions, prompts, memory settings and seeds for every run.

  4. Ablate one component

    Swap the memory store, retrieval policy, summariser or model in isolation.

  5. Repeat runs

    Run each configuration several times to see the spread of outcomes.

  6. Classify failures

    Tag every missed recall point with a cause from the taxonomy.

  7. Report with intervals

    State effects with confidence intervals and threats to validity.

Conceptual flow of the protocol. The ablation and repetition stages loop until every configuration in the matrix has been covered.

Running the memory study step by step

  1. Build a long-horizon task suite

    Write multi-session scenarios modelled on the production failure: facts introduced early, changed midway and needed later. Each carries the ground truth for every recall point, so scoring is mechanical. Include distractors such as similar customers or superseded instructions, because memory fails on near-misses.

    Output
    Scenarios with answer keys
    Owner
    Researchers with domain experts
  2. Make every run reproducible

    Record tool responses once and replay them, so a change in an external API cannot pose as a memory effect. Pin model versions, prompts, sampling settings, memory parameters and seeds, and log them with each run.

    Output
    Replay fixtures and run manifests
  3. Reproduce the failure first

    Run the current production configuration on the suite. If it does not fail the way production does, the suite is missing something and every later comparison is suspect.

    Output
    Baseline failure profile
  4. Ablate one component at a time

    From the baseline, change exactly one element per configuration. Test combinations later, and only for components that showed an effect on their own.

    Output
    Ablation results
  5. Repeat runs and estimate variance

    Agents give different answers to identical inputs. Run each configuration enough times to see the spread, and report differences with confidence intervals rather than single scores.

    Output
    Per-configuration distributions
  6. Classify each failure and report

    Tag every miss using the taxonomy below, then write up effects, intervals, threats to validity and the recommended change.

    Output
    Failure counts by cause and a technical report

An ablation matrix for memory, retrieval and model

Each row changes one thing relative to the baseline. Reading the last column shows what a difference from the baseline would point to.

ConfigurationWhat changesHeld fixedWhat a difference suggests
BaselineNothing; current production designEverythingThe reference profile for every other row
Memory store swapFree-text summaries replaced by structured fact recordsModel, retrieval, promptsLost or contradictory facts trace to how memory is written
Retrieval policy swapSimilarity search replaced by recency-weighted or filtered retrievalModel, store, promptsFacts exist in memory but are not surfaced when needed
Summarisation offRaw transcripts stored instead of summariesModel, store, retrievalDetail is lost during compression
Model swapDifferent model family or sizeStore, retrieval, promptsThe model mishandles context it was given
Oracle memoryCorrect facts injected directly into contextModel and promptsAn upper bound: if this still fails, memory is not the main problem

The oracle row is often the most informative single experiment, because it separates a fact that never arrived from a fact the model could not use.

A failure taxonomy for agent memory and coordination

Lost fact
Information given in an earlier session is absent from memory when it is needed.
Stale fact
A superseded value is recalled, such as an old delivery address.
Contradictory memory
Memory holds conflicting versions of a fact and the agent picks the wrong one or hedges.
Retrieval miss
The correct fact is stored but retrieval does not return it for the current step.
Use failure
The correct fact reaches the model's context but the response ignores or misapplies it.
Coordination failure
In a multi-agent system, one agent holds the fact but never passes it to the agent that needs it, or two agents write conflicting updates.

Threats to validity in agent memory studies

The suite does not resemble production

Early signalThe baseline passes scenarios that fail in production.

MitigationRebuild scenarios from de-identified failure traces before trusting any comparison.

Too few repetitions

Early signalRankings between configurations flip when the study is rerun.

MitigationAdd runs until intervals clearly separate or clearly overlap, and report either outcome.

Hosted model drift

Early signalProvider-side behaviour changes during the study window.

MitigationPin versions where possible and rerun the baseline at the end to detect drift.

Scoring errors

Early signalAutomated scores disagree with human spot checks.

MitigationScore against answer keys where possible and audit a sample of model-graded items by hand.

From finding to fix: the transfer package

The finding is rarely just "change the retriever". A useful package contains the suite and replay fixtures, ablation results with intervals, failure counts by cause, the recommended change with its evidence, and the configurations tried and rejected, because someone will propose those again.

The suite then becomes a regression test: rerunning it after each model upgrade shows whether the fix still holds. Release gates are a separate discipline, covered in our agent production readiness process; this study answers the earlier question of what to build.

Questions and answers

How many scenarios does a useful agent task suite need?

Enough to exercise each failure type several times, with distractors, rather than a fixed count. A small suite built from real failure traces usually teaches more than a large synthetic one. Start with scenarios that reproduce the production problem, confirm the baseline fails on them, then add variations until failure counts by cause stay stable between reruns.

Can we use a public agent benchmark instead of building our own suite?

Public benchmarks help compare models in general, but they rarely contain your domain's facts, distractors or session patterns. Use one as a sanity check alongside a suite built from your own traces. If the two disagree, trust the one that resembles production, and treat the disagreement as a finding worth explaining.

Why record tool responses instead of calling live systems?

Live systems change between runs: prices move, records update and APIs time out. Any of those can alter an agent's behaviour and be mistaken for a memory effect. Replaying recorded responses gives every configuration an identical world, which is the only way a difference can be attributed to the component that changed.

Does this protocol work for multi-agent systems?

Yes, with additions. Each agent's memory and the messages between agents become components that can be ablated, and the coordination category of the taxonomy matters more. Log which agent held each fact and when it was handed on, so a missed recall can be traced to the agent that lost it or the hand-off that never happened.

Sources

  1. Frontier R&D: questions that justify research — ColdAI

More in Frontier R&D

Back to Frontier R&D

Next step

Bring us a context-loss problem you cannot attribute

Describe the failure and share a few de-identified traces. We will say whether a controlled study is warranted and what the first task suite would need to reproduce.

Discuss an agent memory study