How it works
How AI agent memory works
The model does not remember anything. A loop around it does: look up what matters before the answer, write down what was learned after it. Here is one full turn.
-
01
A message arrives
The user asks something. The agent has the current conversation in its context window and nothing else.
-
02
Recall
The question is used to search the user's memory store. A handful of matching notes come back, each a line or two long.
-
03
Assemble the prompt
Instructions, the recalled notes, the last few turns and the new question are put together. This is all the model sees.
-
04
Answer
The model replies. Because the notes are in the prompt, the answer reflects what the user said weeks ago.
-
05
Write
After the turn, or after the session, a second pass decides whether anything new is worth keeping and writes it as a short note.
-
06
Update and forget
If the new note contradicts an old one, the old one is replaced. Notes past their use are removed. The basket stays light.
What the model receives
A prompt with memory is short on purpose
Everything else stays in the store. The size of the prompt no longer grows with the age of the relationship, only with what this question needs.
The five parts of a memory system
Extractor
Decides what in a conversation is worth keeping. Usually a model call with a strict instruction: stable facts only, in the user's own terms, no guesses.
Store
Where notes live. Often a vector index for search by meaning, sometimes with a graph or plain tables alongside for exact facts.
Retriever
Finds the notes that matter for a question. Good ones mix meaning, keywords, recency and importance.
Budget
A hard limit on how many tokens of memory go into a prompt. Without it, memory slowly turns back into a long transcript.
Controls
The user-facing side: see what is remembered, correct it, delete it. Also where isolation between users is enforced.
How it works: common questions
When does an agent write to memory?
Common choices are after every turn, at the end of a session, or when the context window is about to be trimmed. Writing at the end of a session is cheaper; writing before a trim protects against losing facts mid-conversation.
How many memories should be recalled per turn?
Few. Three to ten short notes is a typical range. Set a token budget and let the best matches fill it.
What happens when two memories disagree?
The write step should compare a new note with similar existing ones and replace the older fact, keeping the date. If both are left in place the agent will pick one at random.
Does the model itself change?
No. The model's weights stay the same. Memory lives outside the model and reaches it only as text in the prompt.
Pack the doko. Ask it anything.
Nine memories, one question, no account. See which facts an agent would carry into its next answer.