Guide · 5 min read

How to evaluate AI memory

A practical way to test whether an AI agent's memory works: five kinds of test question, what to measure on each turn, and how to build a small test set.

“It seems to remember” is not a result. Memory fails quietly: the agent still answers, just without the fact it should have had. You need tests that make the absence visible.

Five kinds of question

Write scripted conversations spread over several sessions, then ask:

  1. Single fact. Something stated once, several sessions ago. “What am I allergic to?”
  2. Updated fact. Something that changed. The user said Lisbon, later Porto. Which does the agent use?
  3. Combined facts. An answer needing two memories from different sessions. “Book somewhere for Ana’s birthday that I can eat at.”
  4. Time. “What did we decide last week?” needs dates, not only meaning.
  5. Nothing to recall. A question the memory cannot answer. The right behaviour is to say so, not to invent.

The fifth kind is the one most often left out and the one that catches hallucinated memories.

Measure the two halves separately

A wrong answer has two possible causes, and the fixes differ.

Recall. Was the needed note among those retrieved? If not, the problem is extraction (it was never written) or search (it was not found).

Use. The note was in the prompt, but the answer ignored or contradicted it. That is a prompt or model problem.

Log the retrieved notes for every test question so you can tell which half failed.

What to track

Measure What it tells you
Accuracy on memory questions Whether memory works at all
Needed note retrieved (yes/no) Extraction and search quality
Irrelevant notes retrieved Noise that crowds the prompt
Tokens of memory per turn Cost, and whether the budget holds
Added latency per turn Whether users will feel it
Store size per user over time Whether forgetting works

Always compare with two baselines

  • No memory. If the score barely changes, your questions do not depend on memory.
  • Full history in the prompt. Often the accuracy ceiling for short histories, at a much higher token cost. Memory should come close to it for far fewer tokens.

Keep it honest

Do not tune on the test set until it passes and then report that number. Hold back a few conversations you never look at while tuning. And when you read published benchmark figures, check what was measured, on which data and against which baseline before you compare them with your own.

Quick answers

How do you test if an AI agent remembers?

Run scripted conversations across several sessions, then ask questions whose answers depend on earlier sessions. Score whether the right fact was recalled and whether the answer used it.

What metrics matter for agent memory?

Answer accuracy on memory-dependent questions, whether the needed fact was among those recalled, how many irrelevant facts came with it, tokens added per turn and added latency.

How big should a memory test set be?

Start with twenty to fifty questions that you write by hand from realistic conversations. A small set you trust is worth more than a large one you do not understand.

Read next

Pack the doko. Ask it anything.

Nine memories, one question, no account. See which facts an agent would carry into its next answer.