Context Windows: What the Model Can See in One Run
A context window is the bounded information available to the model for one run. Why that is not memory, what happens when you reach the limit, and why selection still matters at any size.
Ask an assistant something it told you yesterday and it may have no idea what you mean. Ask again in the same conversation and it answers immediately. People reach for the word memory to explain the difference, and that word causes most of the confusion here.
What it actually is
A context window is the bounded information available to the model for the current run.
Everything the model works from occupies that space: the system instruction, the conversation so far, any documents the application retrieved, the current question, and the response being generated as it grows. It is measured in tokens rather than words or characters, which is why token counts and not page counts decide whether something fits.
The important property is the one people skip. The window describes a single run. When the run finishes, nothing in it carries forward on its own.
What it is not
It is not memory.
Persistent memory is a separate application-level mechanism that may store information outside the run and later place selected information into that context. When an assistant remembers your name next week, an application wrote it down somewhere and put it back into the input before the model was called. The model did not retain it. The parameters are unchanged, which is a different page: training, fine-tuning and inference.
- One run
In the window for this run
- The system instruction
- The conversation so far
- Documents the application retrieved
- The current question
- The response as it is generated
Outside the run
- Persistent memory, stored by the application
- Whatever was not selected into this run
- One run
A context window is the bounded information available for one run, and nothing in it carries forward on its own.
This matters practically, not just terminologically. Because that memory lives in ordinary storage, you can inspect it, correct it and delete it. If the model had genuinely absorbed the information, none of those would be available to you.
The phrase "working memory" is a reasonable analogy for the feeling of it, and this page used to lead with it. The analogy misleads on the point that matters most, which is persistence, so it is better to describe the mechanism directly and skip the metaphor.
What happens at the limit
The old version of this page said the oldest tokens are pushed out of memory. That describes what some applications choose to do, not something the model does.
What actually happens when input exceeds the limit depends on who is handling it:
- The model or provider rejects the request, or truncates it according to its own rule.
- The application decides in advance what to leave out, drop, summarise or shorten.
- A chat interface may silently drop or compress older turns on your behalf.
Three different systems, three different behaviours, and only the middle one is under your control. If your system needs specific behaviour at the boundary, the honest answer is that you have to implement it rather than assume it.
Why a bigger window is not the whole answer
Windows have grown enormously, which invites the conclusion that selection stopped mattering. It did not, for two reasons.
The first is cost and speed. Longer input generally means more to process, so it tends to cost more and take longer. How much more depends on the model and the provider, and the only reliable figures are the ones you measure on your own workload.
The second is that filling a window is not the same as using it well. Selection and placement can matter as much as raw context size. You will see one version of this discussed under the name "lost in the middle". This page has not measured it, so treat the general shape as worth testing rather than as a rule: if your task depends on the model attending to one specific instruction inside a long input, verify that it does, on your own task, rather than assuming position is neutral.
What survives from all this is unglamorous. A short, relevant input is easier to get right than a long one, and it is cheaper.
A way to divide the space
This is a PMP working heuristic rather than a standard, and it is useful mainly when a system's inputs have grown messy. Think of the window as holding three kinds of thing:
- Persistent core. Identity, policy, definitions, the constraints that apply to every run.
- Task frame. The current objective and what counts as done.
- Evidence pack. Only the material this particular step needs.
The value is not the three names. It is that the categories have different lifetimes, so labelling them makes it obvious which parts should be identical on every call and which should be rebuilt each time. If your inputs are already tidy, you do not need this.
Finding versus choosing
One distinction is worth carrying between pages.
Retrieval finds candidate material. Context assembly decides what the model actually receives. A search can surface the right passage and the assembly step can still leave it out, usually because something had to go to fit the budget.
If you are working on what gets found, that belongs to retrieval. If you are working on what gets included, you are working on the context window, and this is the page for it.
What this changes
Stop treating the window as a store of what the system knows. It is the space for one run, and what goes into it is chosen, by you or by something you did not configure.
Log the assembled input, not just the retrieved candidates. The difference between those two lists is where a surprising number of "it forgot" reports actually come from.
Where this helps, and where it stops
- What it is
- An explainer of what the context window is, what it is not, and how information is chosen to occupy it.
- What it does not guarantee
- A comparison of specific model context limits or prices, and not a treatment of how candidate material is found, which belongs with retrieval.
- When the distinction matters
- A system behaves as though it forgot something, or a larger window fails to fix a quality problem.