Context Compaction
Context compaction is a technique for keeping a long AI conversation or agent task running when it approaches the model's context window limit. The older part of the history is replaced with a shorter summary, written by the model itself, and work continues from that summary plus the most recent turns. It trades exact detail for space, so long-running work can go on without starting over.
How Context Compaction Works
As a conversation or agent session grows, every message, file read and tool result stays in the context window. When the total gets close to the limit, compaction runs:
- Trigger: compaction starts automatically at a token threshold, or manually when the user or application requests it. Claude Code, for example, offers a /compact command that accepts a focus such as "focus on the API changes."
- Summarize: the model condenses older turns into a summary that aims to keep goals, decisions, key code changes and unresolved issues.
- Replace: the summary takes the place of the original turns. Recent turns may be kept word for word.
- Continue: the task resumes in a much smaller context.
Compaction is often paired with lighter-touch methods. Tools can first clear old tool outputs that are no longer needed, and summarize only if space is still short. Some systems let teams supply their own summarization instructions so specific details are always preserved.
Why Context Compaction Matters
Without compaction, long agent tasks hit a wall: the session fails or must restart from nothing. Compaction also improves quality, not only capacity. Anthropic's documentation notes that response quality degrades as a conversation grows, so keeping the active context small helps the model stay focused.
The cost is information loss. A summary is a judgment about what matters, and it can drop details that turn out to be important later. Claude Code's documentation warns that detailed instructions from early in a conversation may be lost after compaction. That is a common source of agents "forgetting" rules mid-task.
The practical lesson for teams is to never rely on conversation history for anything that must persist. Rules, decisions and requirements belong in durable files, a spec or agent memory, which are reloaded every session regardless of what a summary kept.
Context Compaction Example
A developer spends an afternoon with a coding agent refactoring a payment module. Early on, she tells the agent: "Never log full card numbers, even in debug mode." After many file reads and test runs, the session nears its limit and compacts automatically. The summary records the refactor's progress and the files changed, but not the logging rule.
An hour later, the agent adds a debug log that prints the full payment payload. The fix is to move the rule into the project's instruction file, which the agent loads at the start of every session and after every compaction. For how persistent memory and task-specific retrieval work together so key rules survive, see context management for AI coding agents.
Context Compaction vs. Context Management
Compaction is one tool. Context management is the broader practice of deciding what to store, load, prune and refresh across a task and between sessions. Compaction is a form of the "compress" strategy within context engineering, and works best when the information that must never be lost is kept outside the conversation in the first place.