Guide · 4 min read

Context window vs memory

A larger context window is not memory. What each one does, why long prompts cost more on every turn, and how to decide what stays in context and what moves out.

The context window is what a model can read in one request. Memory is what is still there for the next one. They are often confused because, inside a single chat, a big window feels like memory.

Three limits of the window

It ends with the session. Open a new chat and the window is empty. Whatever the user told the agent yesterday is gone unless something outside the model kept it.

You pay for it every turn. Most chat applications resend the conversation with each new message. Turn 30 carries turns 1 to 29 with it. Over a conversation, total input grows with the square of its length. The context cost calculator shows the curve with your own numbers.

Long prompts are harder to use. Research on long-context models has repeatedly found that information in the middle of a very long prompt is recalled less reliably than information near the start or end. A short prompt with the right five facts is often answered better than a long one that contains them somewhere.

What memory adds

Memory moves facts out of the window into a store, and brings back only those that matter for the current question. The prompt stops growing with the age of the relationship.

Context window Memory
Lifetime One request or session Across sessions
Content Exact recent text Short extracted facts
Cost Grows with every turn Roughly constant per turn
Risk Overflow, lost detail in long prompts Stale or wrong facts if not maintained

Deciding what goes where

Keep in the window: instructions, the last few turns, tool results still in use.

Move to memory: anything the user would be annoyed to repeat next week.

Summarise: the older part of a long conversation, before it is trimmed. Write any durable facts to memory at that moment, because after trimming they are gone.

They work together

Memory does not replace the window; it feeds it. The window is the desk. Memory is the filing cabinet, and recall is the act of taking out the one folder you need.

Quick answers

If context windows keep growing, is memory still needed?

Yes. A window holds one request. It does not carry anything to the next session, it is paid for on every turn, and facts placed deep in a very long prompt are easier for a model to overlook.

What should stay in the context window?

The instructions, the last few turns, and whatever was recalled for this question. Older turns can be summarised or moved to memory.

Does prompt caching remove the cost problem?

It reduces it for the part of the prompt that repeats exactly. It does not help across sessions and does not make a long prompt easier for the model to use.

Read next

Pack the doko. Ask it anything.

Nine memories, one question, no account. See which facts an agent would carry into its next answer.