What Survives Compaction?
When a long LLM conversation gets summarized to fit the context window, which details survive, and why do the rest disappear? I planted facts in long synthetic conversations, compacted them in different ways, and measured what the model could still recall. The short version: a summary loses what it can't know will matter, and the fix is to keep the originals within reach.
TL;DR
- A summary can't know what the next question will be. After a change of plans, OpenAI and Anthropic summaries had kept only 4–13 of the 32 facts the new goal needed. Announcing the next goal in advance restored 30 of 32; naming it just as often without making it next restored 6. Without that knowledge, the only fix at write time is to stop compressing.
- So never let the summary be the only copy. Keeping the original messages and letting a small model pick what to fetch once the question is known brought back 28–32 of the 32. The cheapest model chose as well as the strongest.
- With recall on, the cheapest summarizer is enough. Alone, its summaries keep less than half of what the most thorough one's do; with recall they match it, at about 1/25 of the cost per summary.
- Loss happens in the summary, and it compounds. When the summary dropped a fact, the model couldn't answer. With a cheap summarizer, facts that went through one summary always survived; after four, none did.
- What you ask the summarizer to keep matters as much as which model you use. One instruction change took a cheap model from 31/72 facts to 70/72. And a good summary can beat the raw history: it resolves superseded values that small models otherwise pick.
- Scaffolding has to keep earning its place. Every mechanism between the conversation and the model is re-measured against going without it. The first audit caught recall misleading a weak model; after the summarizer changed, the second found recall essential.
- Watch the bill. Reasoning tokens were the biggest cost, my early estimates were off by 2x, and a timestamp at the top of the prompt had silently disabled prompt caching.
The question
Every long-running LLM app eventually hits the context window. The common fix is compaction: summarize the older history, keep going from the summary. Claude Code does it, Codex CLI does it, and so does The Crucible, an open-source multi-model chat app I maintain.
My question was simple: is there a point where the model suddenly forgets? And if details do get lost, why? Is it the summary, the model's attention, or something about doing it repeatedly?
I couldn't find a controlled answer, so I built one into the app. Everything below runs on the app's real LangGraph pipeline, and the code and raw results are public.
How I measured it
Planted facts. Each synthetic conversation is 140 exchanges, about 23k tokens, with 24 fictional facts buried mid-message: budget caps, reviewers, key labels, ports, dates, quotes, vendor choices, and "we agreed not to use X". Fictional values mean the model can't answer from prior knowledge.
Make details compete. My first pilot scored 100% everywhere because the facts were the only specific content in generic filler. So every turn now carries specifics of the same shape (latencies, ticket counts, owners), and facts come in three variants:
- Plain: stated once.
- Distractor: a rejected proposal with a different value sits nearby.
- Update: an older value appears first and is later changed. Answering with the old or rejected value counts as wrong.
Compact the way the app does. The conversation is replayed turn by turn through the app's own summarize step, so a long conversation gets summarized several times as it grows. Then each fact is probed with a question such as "Which port does the CDN purge job listen on? Answer with just the value."
Separate the summary from the model. An oracle pseudo-model answers correctly if and only if the value is still in the context it was given. Its recall measures what the summary kept. A real model scoring below it means the information was there but missed. A stricter oracle also requires the value to still be attached to its subject.
Judge in pairs. The same fact is probed under each condition, and conditions are compared with a paired proportion test and a 10-point equivalence margin. I used my own Ordal library for this.
Cap the spend. After an expensive surprise (see below), every run first makes one calibration call per model, dry-runs the whole design for free to count tokens, and refuses to start if the estimate exceeds a hard budget.
Finding 1: a good summary can beat the raw history
With no compaction at all, three small models missed 12–19% of facts in a 23k-token conversation. Most misses were updated facts: gemini-3.5-flash-lite gave the superseded value for 10 of 18 of them.
Compacting with a strong summarizer fixed that. The summary kept every fact and resolved each update to its latest value, so stale answers dropped to zero.

A summary is not just a shorter transcript. It is a consolidated state, and that can be easier for a model to use than the raw history.
Finding 2: loss compounds with every pass
With a cheap summarizer, the number of summaries a fact has been through predicts whether it survives. After four passes, nothing was left.

Is it the passes, or a lack of room? Two explanations fit that curve. Either each pass drops a little of what the previous one kept, or a fixed-size summary pushes out the oldest facts. Because older facts have always been through more passes, the two are confounded.
So I summarized the same conversations two ways with the same model: incrementally as they grew, or once at the end.

A single pass kept most of the early facts that incremental summarizing lost. That points to repeated passes rather than lack of room for old facts, though the overall paired test was not decisive (+12.5 points, 95% interval −3.5 to +27.7).
One big pass has its own cost, though. It kept 92% of values somewhere, but only 75% still attached to the right subject. It loses who a value belonged to.
Repetition protects a fact. I restated each fact once, later in the conversation, changing nothing else. Recall rose from 23/72 to 41/72, +25 points (95% interval +12 to +37). Part of that gain is that the restatement is newer, so it went through fewer passes.
Finding 3: the instruction matters as much as the model
The app's original summary prompt asks for "a concise, detailed technical brief" that preserves who argued what. I compared it with Codex CLI's handoff prompt and with a prompt I wrote from first principles, "state and index". That prompt asks for every decision, agreed value and open item with who stated it, only the latest value of anything changed, and a one-line index of topics.
Half the facts in these conversations were stated by the assistant rather than the user.

Telling a cheap model what kind of thing must survive took it from 31 to 70 facts. It paid for that by writing twice as much. A capable model kept nearly everything even in Codex's roughly 1,300-token handoffs, so the limit is what the summarizer chooses to keep, not the room it has.
What a summary costs. The app's old default summarizer, gemini-3.5-flash, spent 10–20k reasoning tokens per summary at $9 per million output tokens: about $0.21 per summary. gemini-3.8-flash kept the same facts (equivalent within 10 points) for about $0.03, so the app switched to it, until Finding 7 below. Cutting 3.5-flash's thinking to "medium" matched the default; "minimal" cost a quarter but lost 18 points.
A bug surfaced along the way. The app counted the summary itself as new material, so once a summary outgrew the threshold, it re-summarized every few messages: 132 summaries where about 20 were needed. The fix requires half a threshold of genuinely new text before summarizing again.
How Codex CLI and Claude Code do it
Codex CLI is open source, so I read its compaction code. Claude Code isn't, so for it I relied on the official docs.

Codex even warns after compacting that "long threads and multiple compactions can cause the model to be less accurate." That is Finding 2, stated by its authors.
I added Codex's rules to the experiment. Keeping user messages verbatim rescued exactly the user-stated facts: with a cheap summarizer, they went from 17/36 to 36/36, while assistant-stated facts did not move. The cost was 2.3 times more context on every later call.
The principle: never let the summary be the only copy
It is tempting to say coding agents and conversational apps simply need different compaction. I think the real distinction is deeper. Every piece of context has three properties:
- Authority: who said it. A user's statement is a requirement; the assistant's is a claim; a tool's is an observation.
- Recoverability: can it be fetched again? A file can be re-read; a detail said only in chat cannot.
- Liveness: will it be needed? Current values, decisions and open items will be. Superseded values and dead ends won't.
Read that way, Codex's design makes sense. User messages are authoritative and unrecoverable, so they stay verbatim. Tool output is recoverable from the repository, so it can go. The summary only has to carry live state, which is why a short handoff is enough. Claude Code's CLAUDE.md does the same thing by moving authoritative rules out of the conversation, where compaction can't touch them.
Every loss I measured was information whose only copy was the summary. The Crucible never deletes history; it keeps every message in its checkpoint store. The model just couldn't see it after pruning. So the fix is not a better summary but making the original reachable again.
Finding 4: bringing the original back rescues a lossy summary
I added a recall step. Before each call, it searches the hidden history for the four messages that best match the latest request and quotes them back into context. The search is plain keyword matching weighted by word rarity.

Recall added 50 and 79 points to the lossy configurations, and changed nothing where the summary already kept everything. The cheapest setup that kept nearly every detail was the cheapest summarizer writing 1k-token handoffs plus recall: 66/72, with each later call reading 2.8k tokens.
Once the original is reachable, the summary no longer has to be a complete record. It only has to be good enough to continue from.
Finding 5: recall moves the relevance decision to the retriever
Recall only worked so well because my questions named their subject. So I rewrote every question to describe the subject instead: "Which network port should the firewall allow for the task that clears cached files at the edge?" rather than "Which port does the CDN purge job listen on?" No subject word appears in any of them.

When the facts are in view, wording doesn't matter: the model maps "the task that clears cached files at the edge" to the CDN purge job almost every time. Recall hands that judgment to the retriever, and a keyword or small-embedding retriever is far worse at it than the model. Letting a cheap model read the summary and choose the search terms recovered most of the loss, but not all.
So both designs fail on capability, just in different places. A summary fails in the summarizer, compounding with every pass. Recall fails in the retriever. Whatever decides relevance has to understand the question about as well as the model does.
A prediction that failed
I expected facts that sound unimportant when they are said to be dropped more often, even by strong summarizers. So I restated every fact as a throwaway remark ("Side note, probably irrelevant: the staging database listens on port 6543"), changing nothing else.
The opposite happened. A cheap summarizer kept more of the asides: 18/72 instead of 1/72 with Codex-style handoffs. The strong summarizer and recall kept everything either way. In filler where every sentence has the same shape, a hedge makes a sentence stand out rather than fade. Sounding unimportant is not the same as being unimportant to a model, and the harder question, relevance nobody could have foreseen when the fact was stated, remains open.
Finding 6: summaries drop what the task doesn't need yet
So I tied relevance to the task instead of the wording. Each conversation pursues one stated goal ("this week's only priority is the log shipper"). Eight facts are about that goal. Eight more are about another subsystem nobody is working on yet, stated plainly in passing, in exactly the same form as everything else, and mentioned nowhere else. The last turn switches the goal: "Change of plans: the warehouse loader is now the priority." Then every fact is probed.
Until this point every summarizer I had tested was a Gemini model, so I ran two from each provider family on four conversations:

Four of the four OpenAI and Anthropic summarizers lose off-goal facts, and the OpenAI ones keep almost none. Gemini, the family I had been testing all along, does it least.
My first reading, from two Gemini conversations, was that stronger summarizers drop more. It doesn't hold across families: the stronger model is more selective within OpenAI and less selective within Anthropic. The family's summarizing style matters more than strength. What does hold is the mechanism: a task-focused summary leaves out what the task doesn't need yet, and no summarizer can know the plan is about to change.
Recall brings them all back. Recall looks at the history after the goal has switched, so it should not care about the change of plans. I tested it with the two most goal-selective summarizers:

Deciding at read time recovers exactly what deciding at write time loses. This is the one result in the study I expect to survive new models: a summary is written for the goal of the moment, no summarizer can anticipate a change of plans, and only a design that keeps the originals reachable is immune. One caveat carries over from Finding 5: these questions named their subject, so keyword recall was enough. With the same questions asked indirectly, keyword recall restored only 16–17 of the 32 and embedding recall did no better, while model-guided recall, where a small model reads what is still visible (including the turn that changed the goal) and chooses what to fetch, restored 28–30. The two summarizers here are two replications, not a comparison: they differ in family and tier.
Choosing what to fetch doesn't need a strong model. Until then the model choosing the search terms had always been gpt-6-luna, the cheapest in the app. So I fixed one summarizer (gpt-6-sol, the most goal-selective), wrote its summaries once, and replayed them with six different models choosing what to fetch, a cheap and a strong one from each family. Same indirect questions, four conversations:

Every model beats keyword recall. Luna and sol are statistically equivalent, and the strong models are equivalent to luna; only the smallest Gemini falls somewhat behind. The cheapest model made the call as well as gemini-3.8-flash, which cost 24 times as much for the same calls ($0.59 against $0.024). That fits the time-of-decision story: before the question is known, no summarizer can judge relevance reliably; once it is known, even a small model can. Which cheap model is good enough will change with every release. That deciding late makes the decision easy should not.
Why the summary drops them. "No summarizer can anticipate a change of plans" is a claim, so I tested it. I reran the same conversations with one change: every goal reminder also said "once that is done, the next priority will be X". The facts, filler and positions stayed identical. Separately, I tried a generic hedge that knows nothing about what comes next: the usual instruction plus "priorities may change; keep the specific values for every subject, not only the current goal."

For gpt-6-sol, knowing the next goal is the whole story: 4 to 30 of 32, for summaries 30% longer. The generic hedge also kept everything, but only by barely compressing: summaries of 5.5k tokens against an 8k threshold, rewritten twice as often. Without knowing the future, the only way to keep off-goal facts at write time is to give up the compression. Claude Haiku shows a second, separate loss: told the future and hedged, it still kept only 20 of 32, and it lost on-goal facts too. That looks like fidelity, not selection. Read-time recall sidesteps both, because the originals are still there. Was it only that the subject's name kept coming up? Announcing the next goal names it a dozen times, so I ran a control that names it just as often, as out of scope this quarter, with the change of plans still a surprise. gpt-6-sol kept 6 of 32: about 2 of the 26 recovered facts come from the name standing out, and the other 24 from knowing it comes next.
Finding 7: once the originals are reachable, the cheapest summarizer is enough
Recall changes what a summarizer is for, so I chose the app's default summarizer again, this time with recall on, as the app runs. Before running, I wrote down the rule: keep the candidates statistically no worse than the best on both scenarios, then take the cheapest per summary at the prices in force from January 2027 (gemini-3.8-flash's introductory price ends in December). Four summarizers, the detail-heavy scenario and the change-of-plans scenario, four conversations each, strict oracle:

Only gpt-6-luna passed the rule: statistically equivalent to the best on the detail-heavy scenario, itself the best after a change of plans, and about 1/25 of gemini-3.8-flash's cost per summary (costs are for the detail-heavy scenario at 2027 prices). Alone, its summaries keep less than half as much. Once the originals can be fetched, the summary's job shrinks from being the record to being a pointer, and a pointer can be cheap. The choice holds only because recall is on: without it, the ranking inverts.
One result I can't explain yet: recall added 16 to 29 facts for every other summarizer, but nothing for gemini-3.8-flash after the change of plans (56 either way), and a later audit reproduced the gap.
Keeping the scaffolding honest
Every mechanism in this post, from the summary itself to the verbatim window and recall, exists because it beat a baseline on the models I tested. A newer model can close that gap on its own, and scaffolding that no longer helps can get in its way. So each mechanism on the app's default path now declares its baseline and its claim, and an audit re-measures the pair on current models whenever the model list changes, recording keep, retire, harmful or inconclusive.
The first two audits already changed the picture. Both used two cheap answering models on the detail-heavy scenario:

With a thorough summary, the recalled excerpts were a distraction for the weaker model: it took values from the wrong subject and gave the same answer to different questions. With a thin summary, the excerpts are the main source, and recall became essential for both. The second audit also checked the summarizer choice with real answering models: neither pricier summarizer did better for either reader.
Not every verdict should be acted on. The audit flagged the verbatim window, which keeps the newest messages as they are, as "retire": with recall on, it added nothing to fact recall. But the probes only ask about planted facts, and the window exists for continuity with the last few turns, which they don't measure. An audit is only as good as its claim.
What changed in the app
- The re-summarize loop is fixed, and prompt caching works on all three providers.
- Recall is on by default: once a thread has been summarized, the cheapest available model picks which original messages the new request is about and shows them again. It costs nothing until a thread compacts, then one small call per turn.
- The default summarizer is gpt-6-luna, chosen above, and any other can be picked in the Control Panel.
- Every mechanism on the default path is re-audited when the models change.
- State-and-index summaries, keyword recall and keeping user messages verbatim are measured here but are not defaults.
What it cost, and what went wrong
The whole study cost roughly $80 in API calls. About $41 of that went on the first four runs, before I had a budget guard. The rest, about twenty more experiments, cost about $36 under hard caps, plus about $3 that a cost estimate spent by mistake, including one run that returned nothing and a few dollars of calibration for runs that were then refused.
The mistakes were as instructive as the results:
- A ceiling effect. The first pilot scored 100% everywhere because the planted facts were the only specific content in the conversation.
- A scoring bug. Models often answer and then quote their source, such as "Litware. It said: 'we picked Litware over Margie and Contoso'". Matching anywhere in the reply counted the quoted rejected vendors as wrong. Scoring now reads only the committed first line.
- The re-summarize loop from Finding 3, found only because the first expensive run wrote 132 summaries.
- Guessed prices. I assumed a summarizer cost $2.50 per million output tokens; it was $9. I had also picked "nano" models as the cheap default by name, while a newer model cost half as much. Prices now live in a table with a source for each, and experiment defaults are chosen from it.
- A strict oracle that was too strict. My first "value next to its subject" check used a 200-character window. Summaries organized in long per-topic sections broke it, while every model still answered correctly.
- A noisy estimator. One calibration call per model is not enough. One run cost 1.8 times its estimate. The hard cap, not the estimate, is what bounds the spend.
A capped run that returned nothing. A long-horizon run (about 22 summaries per conversation) spent its whole $5 cap and produced zero results. Every condition was replaying at once, the summaries grew with every fold faster than my one-call calibration predicted, and the cap stopped all of them mid-way. Runs now go one condition at a time, cheapest first, each admitted only if the remaining budget covers it and saved as soon as it finishes.
An estimate that spent money. The cost estimate is a dry run with stand-in models, so it should be free. But the model choosing what to fetch was handed the real model client when the experiment was set up, and the dry run never swapped it out. Estimating the six-model run above made about 384 real calls, roughly $3, recorded nowhere and priced at zero. Now each step that calls a model says which model it calls, and the estimate meters it with a stand-in.
The cache that never hit. Cache hits across the experiments were about 1%. The app's system prompt began with the current time to the second, so no two requests shared a prefix. Moving the time out of the system prompt was not enough, because each provider caches differently:

The general rule: anything that changes per request (a timestamp, retrieved excerpts) belongs after the last point the next request will repeat, and where that point is depends on the provider.
Related work
These results sit next to a growing body of work, and several papers point the same way:
- The Complexity Trap (JetBrains Research, NeurIPS 2025 DL4Code workshop): in coding agents, simply masking old tool output matches LLM summarization at about half the cost.
- Evaluating Context Compression (Factory.ai): probe-based evaluation of compaction in real agent sessions. It credits merging each summary into a persistent state instead of regenerating it, which lines up with the state-and-index result here.
- CliffCompaction (Nguyen, Cho, Chen, Dettmers, 2026): never compacts already-compacted content, to avoid drift. That is the per-pass loss measured in Finding 2.
- Toward Reliable Context Compression for Long-Horizon Agents (2026): repeated compression makes agent runs unstable.
- What Does Context Compression Cost an Agent? (2026): completion rates stay flat while agents re-fetch dropped state, so the cost hides in extra calls.
- Fidelity Before Structure (2026): for conversation memory, retrieving verbatim chunks beats retrieving extracted artifacts.
- Background: recursive summarization for dialogue memory, Letta/MemGPT's tiered memory with searchable recall, Anthropic's context engineering guide, and Chroma's Context Rot.
What I haven't seen elsewhere, though I may have missed it: a controlled once-versus-incremental comparison to separate per-pass loss from capacity, an oracle that separates "the summary dropped it" from "the model missed it", a paired test of repetition, a causal test of why task-focused summaries drop facts (announcing the next goal versus only naming it), and an audit that re-checks each piece of context machinery against going without it whenever the models change.
Limitations and what's next
This is a pilot study, not a benchmark:
- Synthetic conversations. Real conversations are messier, and the next step is to replay real threads.
- Small samples. Each cell is 32 to 96 paired probes from three or four conversations, with small and mid-size models only.
- One threshold and no multiple-comparison correction. Treat single comparisons as indicative.
- Task relevance is preliminary. Findings 6 and 7 rest on four conversations per condition, scored by the oracle. Real answering models were checked only on the detail-heavy scenario, in the audits, not after a change of plans. Real conversations are still to come.
- String-match scoring. It is strict, and it only reads a reply's first line.
Next I want to run it on real conversation histories, and add stronger models. The goal is to turn this into a small open-source context-management layer that decides what to keep verbatim, what to summarize and what to leave retrievable.
Code and data. Everything runs in The Crucible (crucible compaction-eval and crucible scaffold-audit). Every number here is in the lab notes, with the raw results next to them.
How this was made. I ran this study with Claude Code as a coding and analysis assistant. It wrote much of the experiment code and the first draft of this post. I directed the questions and checked the results, and I'm responsible for the conclusions.