Mitigation
Context Compaction Failure Modes: What Summarization Drops and How to Catch It
Context compaction failure modes are the specific ways an agent's context-summarization step silently discards information the agent still needs: recency-biased eviction, over-aggressive precision tuning, and fidelity loss compounding across repeated summarize-and-restart cycles. A breakdown of the mechanism, the three concrete failure patterns, and detection and mitigation steps grounded in Anthropic's and MemGPT's published compaction research.

Context compaction failure modes are the specific, recurring ways a summarize-and-restart step designed to keep an agent's context window usable ends up discarding information the agent still needs, three or ten turns later, with no error and no visible signal that anything was lost. Compaction itself is not the failure: it is the standard fix for a context window approaching its limit, and Anthropic describes it plainly as "the practice of taking a conversation nearing the context window limit, summarizing its contents, and reinitiating a new context window with the summary" (Anthropic, "Effective context engineering for AI agents"). The failure is what happens when that summarization step is tuned wrong, run too many times, or trusted without a check: a constraint, a rejected approach, or a half-finished decision gets compressed away, and the agent proceeds as if it had never existed.
This is a narrower, mechanism-specific companion to context rot, which describes reliability degrading as a window simply grows long. Compaction is the standard countermeasure to context rot, which means compaction failure is the case where the fix itself becomes the new source of the same symptom: a smaller, cleaner-looking window that is nonetheless missing something load-bearing.
The mechanism
Compaction runs on a trigger, usually a token-count threshold, and produces a summary that replaces the raw transcript it was generated from. Anthropic's own account of building this into Claude Code describes tuning it in two passes: first maximizing recall so nothing important is dropped, then iterating on precision to trim what is unneeded, using complex agent traces as the test bed (Anthropic, "Effective context engineering for AI agents"). Their implementation preserves architectural decisions, unresolved bugs, and implementation details, discards stale tool calls and redundant outputs, and additionally keeps the five most recently accessed files alongside the compressed summary, rather than trusting the summary alone to carry that detail.
MemGPT's earlier formulation of the same problem names the trigger points explicitly: a warning threshold at roughly 70% of context capacity inserts a system alert giving the model a chance to proactively save what it judges important, and a flush threshold at 100% capacity evicts around half of the queued messages, generating a new recursive summary from the old summary plus the evicted content (Packer et al., "MemGPT: Towards LLMs as Operating Systems," arXiv:2310.08560). Critically, MemGPT does not treat the summary as the sole surviving record: evicted messages stay in a separate recall store, retrievable on demand, which is precisely the property that a bare compact-and-discard implementation lacks. That difference is the hinge this whole topic turns on: compaction that replaces raw history with only a summary has no fallback when the summary is wrong; compaction that keeps the discarded material retrievable does.
Three failure patterns
**Recency-biased eviction.** Flush-based compaction removes the oldest queued content first, which is a reasonable default and also a specific bias: a constraint stated at the very start of a session, a requirement the user mentioned once and never repeated, is exactly the content most likely to be evicted first and therefore most likely to survive only if the summarization step explicitly re-surfaces it. This is the same asymmetry documented in general long-context research, where model performance on retrieving relevant information is strongest when that information sits at the beginning or end of the input and drops significantly when it sits in the middle of a long context (Liu et al., "Lost in the Middle: How Language Models Use Long Contexts," arXiv:2307.03172). A summary is itself a context the model has to read back from later, so information that lands in the summary's own middle, buried under more recent, more salient detail, inherits the same retrieval weakness the original long-context problem had, just at a smaller scale and one level removed.
**Over-aggressive precision tuning.** Anthropic's own description of the risk is precise: the failure surfaces when compaction is tuned toward precision, cutting content judged unnecessary, and ends up cutting "subtle but critical context whose importance only becomes apparent later" (Anthropic, "Effective context engineering for AI agents"). This is a tuning failure, not an architectural one: a compaction prompt that has been trimmed to keep summaries short and clean is optimizing for the wrong thing if the metric being watched is summary length or readability rather than downstream task success several turns later. The detail that gets cut is, by construction, the one that did not look important at the moment of compaction. That is what makes it dangerous rather than merely lossy.
**Compounding fidelity loss across repeated cycles.** A session that compacts once has one lossy transformation between it and the original transcript. A long-running session that compacts five or ten times, each new summary built from the prior summary plus new content, accumulates that loss recursively, and nothing about the mechanism bounds how much drifts away over many cycles. MemGPT's own recursive-summary design (generating each new summary from "the existing recursive summary and evicted messages" rather than from raw history) is explicit about this compounding structure, and while its evaluation shows the underlying retrieval mechanism performs well on tasks like Deep Memory Retrieval, that result depends on the separate recall and archival stores staying available as ground truth to correct against, not on the recursive summary being trusted as a complete record on its own (Packer et al., arXiv:2310.08560). A compaction implementation without an equivalent uncompacted fallback has no such correction available; each cycle's errors simply become the next cycle's starting assumptions.
Why this differs from adjacent failure modes
Context poisoning is false content entering the window and being treated as fact; a compaction failure drops true content instead, which tends to produce a silent gap rather than a confidently wrong claim. Context handoff loss is the same shape of failure, information failing to survive a transition, but at the boundary between two separate agents' contexts rather than within one agent's own compaction cycle; the detection and mitigation patterns below are close cousins of that article's for exactly this reason. Redundant re-exploration is frequently the visible symptom of a compaction failure: the agent re-does work it already completed because the record of having done it did not survive the summary, which is a relatively benign outcome compared to the agent proceeding on a stale or absent constraint without redoing anything at all.
Detecting it
- Log the pre-compaction transcript and the post-compaction summary as separate, diffable artifacts, not just the summary. Without the original, a suspected loss can never be confirmed, only guessed at.
- Seed a session with a small number of specific, checkable constraints early on, then verify after each compaction cycle that all of them are still present or correctly referenced in the summary. A constraint that silently disappears after cycle two but not cycle one localizes exactly which compaction pass caused the loss.
- Watch for redundant re-exploration as a downstream signal: an agent re-deriving a fact or re-running a check it already completed is frequently evidence that compaction dropped the record of having completed it, even when no other symptom is visible.
- Track compaction cycle count per session alongside task accuracy. If accuracy on constraint-adherence checks degrades as cycle count rises even though the summary length stays flat, that is the compounding-fidelity pattern, not a one-off tuning miss.
- Treat tool result clearing as a lower-risk, separately tunable step. Anthropic notes it as "one of the safest, lightest-touch forms of compaction" precisely because stale tool outputs are rarely load-bearing the way a decision or constraint is, which makes it a reasonable first lever to pull before touching decision-bearing content (Anthropic, "Effective context engineering for AI agents").
Mitigation
Keep the pre-compaction record retrievable rather than discarding it outright, mirroring MemGPT's recall-store design: the compacted window is what the model reasons over turn to turn, but a query path back to uncompacted history means a suspected gap can be checked and repaired instead of accepted as unrecoverable.
Move anything genuinely load-bearing, the current goal, hard constraints, open decisions, out of the compaction path entirely rather than trusting it to survive summarization. This is the same discipline state externalization argues for generally, applied specifically to the content most likely to be silently dropped by a lossy summarization step: a goal anchor or constraint file that compaction never touches cannot be a compaction casualty.
Re-anchor explicitly immediately after every compaction cycle, not only on a session-length timer. Re-anchoring patterns are usually framed around long single-agent sessions generally; a compaction event is a specific, detectable moment where re-injecting the externalized goal and constraint set is cheap and directly targets the failure window this article describes.
Tune for recall first, precision second, and re-validate periodically. Anthropic's own two-pass tuning order, maximize recall, then trim for precision, is worth treating as a standing discipline rather than a one-time setup step, because a compaction prompt tuned against last quarter's traces can quietly drift toward over-trimming as task patterns change, with no alert unless someone is specifically checking constraint survival.
Pair compaction with drift regression tests that specifically exercise multi-cycle sessions. A single compaction event is comparatively easy to audit by eye; the compounding failure mode only shows up at scale, across many cycles, which is exactly the kind of regression a one-off manual check will miss.
Related terms
- Context rot: the general degradation-with-length problem that compaction exists to counter, and whose fix is this article's specific failure surface.
- Context poisoning: the adversarial counterpart, false content entering context, versus true content silently leaving it.
- Context handoff loss: the same completeness failure at an agent-to-agent boundary rather than within one agent's own compaction cycle.
- Redundant re-exploration: the common downstream symptom when compaction has dropped the record of work already done.
- State externalization: the mitigation pattern this article applies specifically to compaction-vulnerable content.
- Re-anchoring patterns: the recovery mechanism this article recommends triggering on every compaction event.
FAQ
Is compaction itself a bad practice? No. It is the standard, necessary response to a context window approaching its limit, and both Anthropic's and MemGPT's published designs treat it as the default mechanism rather than an edge case to avoid. The failure modes here are about specific ways the mechanism is implemented or tuned, not an argument against compacting at all.
How is this different from just running out of context? Running out of context is the problem compaction solves. This article is about the new, narrower failure surface compaction itself introduces once it is in place: information that survives the token-limit problem but does not survive the summarization step meant to fix it.
Does keeping a full transcript around solve this? It solves the recoverability half of the problem, matching MemGPT's recall-store approach, but not the detection half. A retrievable original is only useful if something prompts a check against it; pairing retrievability with the constraint-survival checks and drift regression tests described above is what turns a possible recovery into an actual one.
Is there a safe amount of compaction? The evidence points to tool result clearing being comparatively low-risk and decision- or constraint-bearing content being comparatively high-risk, per Anthropic's own guidance. There is no single safe threshold across all content types; the mitigation is treating those categories differently rather than compacting everything with one uniform pass.