How it works

How AI agent memory works

The model does not remember anything. A loop around it does: look up what matters before the answer, write down what was learned after it. Here is one full turn.

  1. 01

    A message arrives

    The user asks something. The agent has the current conversation in its context window and nothing else.

  2. 02

    Recall

    The question is used to search the user's memory store. A handful of matching notes come back, each a line or two long.

  3. 03

    Assemble the prompt

    Instructions, the recalled notes, the last few turns and the new question are put together. This is all the model sees.

  4. 04

    Answer

    The model replies. Because the notes are in the prompt, the answer reflects what the user said weeks ago.

  5. 05

    Write

    After the turn, or after the session, a second pass decides whether anything new is worth keeping and writes it as a short note.

  6. 06

    Update and forget

    If the new note contradicts an old one, the old one is replaced. Notes past their use are removed. The basket stays light.

What the model receives

A prompt with memory is short on purpose

Everything else stays in the store. The size of the prompt no longer grows with the age of the relationship, only with what this question needs.

The five parts of a memory system

Extractor

Decides what in a conversation is worth keeping. Usually a model call with a strict instruction: stable facts only, in the user's own terms, no guesses.

Store

Where notes live. Often a vector index for search by meaning, sometimes with a graph or plain tables alongside for exact facts.

Retriever

Finds the notes that matter for a question. Good ones mix meaning, keywords, recency and importance.

Budget

A hard limit on how many tokens of memory go into a prompt. Without it, memory slowly turns back into a long transcript.

Controls

The user-facing side: see what is remembered, correct it, delete it. Also where isolation between users is enforced.

How it works: common questions

When does an agent write to memory?

Common choices are after every turn, at the end of a session, or when the context window is about to be trimmed. Writing at the end of a session is cheaper; writing before a trim protects against losing facts mid-conversation.

How many memories should be recalled per turn?

Few. Three to ten short notes is a typical range. Set a token budget and let the best matches fill it.

What happens when two memories disagree?

The write step should compare a new note with similar existing ones and replace the older fact, keeping the date. If both are left in place the agent will pick one at random.

Does the model itself change?

No. The model's weights stay the same. Memory lives outside the model and reaches it only as text in the prompt.

Pack the doko. Ask it anything.

Nine memories, one question, no account. See which facts an agent would carry into its next answer.