Context Window
A context window is the total amount of information a large language model (LLM) can take into account in a single request, measured in tokens (small chunks of text, roughly a short word or part of a word each). It works like the model's working memory: everything the model reads and writes for one response, including instructions, conversation history, files and its own output, must fit inside it.
How a Context Window Works
Every request to a model is assembled from several parts, and all of them count toward the window:
- System instructions: standing rules for how the model should behave.
- Conversation history: earlier user messages and model replies.
- Supplied material: documents, code files, images and search results.
- Tool definitions and results: descriptions of available tools and the output they return.
- The response itself: including any reasoning tokens the model generates.
In a multi-turn conversation, each turn is added to the history, so the window fills up over time. If the input alone exceeds the limit, the request fails. Long-running tools and agents therefore manage the window actively, for example by clearing old tool output or summarizing earlier turns through context compaction.
The window is separate from the model's training data. A model may "know" general facts from training, but it only knows your specific product, codebase or decisions if that information is placed in the window for the current request.
Why the Context Window Matters
The window sets a hard limit on what an AI can reason about in one step. If a requirement, a business rule or an earlier decision is not in the window, the model cannot use it and will often fill the gap with a plausible guess.
Bigger is not automatically better. Anthropic's documentation notes that as the token count grows, accuracy and recall degrade, an effect it calls context rot. A window stuffed with loosely related material can produce worse answers than a smaller window holding exactly the right facts. Large windows also cost more, since you pay for every token sent.
This is why the practical question for product and engineering teams is rarely "how large is the window?" It is "what belongs in it for this task?" That selection work is the focus of context engineering.
Context Window Example
A team asks an AI coding agent to add a discount code field to checkout. The request includes the agent's instructions, the story and acceptance criteria, five relevant source files and the test output. Together they use a modest share of the model's window, and the agent completes the task.
Two hours later, after dozens of file reads and test runs in the same session, the window is close to full. The agent's tool automatically summarizes older turns to free space. The summary keeps the main code changes but drops a detail from early in the session: discount codes must not stack with annual pricing. The next change violates that rule. Writing such rules into a persistent project file, rather than relying on conversation history, avoids the loss. Why a large window does not remove the need to choose the right information is explained in context management for AI coding agents.
Context Window vs. Agent Memory
The context window is temporary and per request. Agent memory is information stored outside the window, such as notes or files, that can be loaded back in when needed. Memory lets work survive across sessions; the window determines what the model can see right now.