Calculator

LLM context cost calculator

Resending the whole conversation on every turn is the simplest design and the one that costs most as chats get longer. Compare it with recent turns plus recalled memory, using your own numbers.

An example figure. Enter your own model's price, in your currency.

Input cost per month

Send the full history every turn

tokens per conversation · longest prompt

Recent turns plus recalled memory

tokens per conversation · longest prompt

The rules behind the numbers

Let T be the tokens added per turn and N the number of turns.

  • Full history. Turn t sends t × T tokens. Over the conversation that is T × N × (N + 1) ÷ 2.
  • With memory. Turn t sends the last k turns plus M tokens of recalled memory: min(t, k) × T + M.
  • Monthly cost. Tokens per conversation × conversations per month × price per million ÷ 1,000,000.

This is a simple model. It ignores output tokens, the system prompt, prompt caching and the cost of writing memories. It shows how the two designs scale, not what your invoice will say.

For the reasoning behind the two designs, read context window vs memory.

About the calculator

Why does sending the full chat history get expensive?

Each turn resends everything before it. The total input over a conversation therefore grows with the square of its length: twice as many turns is about four times the input tokens.

How does memory reduce input tokens?

The prompt holds only the last few turns and a small set of recalled notes, so its size stays roughly constant however long the conversation runs.

What does this calculator leave out?

Output tokens, the fixed system prompt, the model calls used to write memories, storage, and prompt caching, which many providers offer and which lowers the price of repeated history. Treat the result as a comparison of shape, not a quote.

Which price should I enter?

The input price per million tokens published by your model provider. The default is only an example so that the bars have something to show.

Is memory always cheaper?

No. For short conversations the full history is small and memory adds its own tokens. Memory pays off with long conversations and with users who come back.

Numbers done. Now see recall.

The playground shows which facts an agent would carry into a prompt, and which it would leave behind.