← 回到 The CrucibleBack to The Crucible

,閱讀約 20 分鐘, 21 min read

這篇原文是英文,這裡是中文翻譯。

What Survives Compaction?

When a long LLM conversation gets summarized to fit the context window, which details survive, and why do the rest disappear? I planted facts in long synthetic conversations, compacted them in different ways, and measured what the model could still recall. The short version: a summary loses what it can't know will matter, and the fix is to keep the originals within reach.

TL;DR

The question

Every long-running LLM app eventually hits the context window. The common fix is compaction: summarize the older history, keep going from the summary. Claude Code does it, Codex CLI does it, and so does The Crucible, an open-source multi-model chat app I maintain.

My question was simple: is there a point where the model suddenly forgets? And if details do get lost, why? Is it the summary, the model's attention, or something about doing it repeatedly?

I couldn't find a controlled answer, so I built one into the app. Everything below runs on the app's real LangGraph pipeline, and the code and raw results are public.

How I measured it

Planted facts. Each synthetic conversation is 140 exchanges, about 23k tokens, with 24 fictional facts buried mid-message: budget caps, reviewers, key labels, ports, dates, quotes, vendor choices, and "we agreed not to use X". Fictional values mean the model can't answer from prior knowledge.

Make details compete. My first pilot scored 100% everywhere because the facts were the only specific content in generic filler. So every turn now carries specifics of the same shape (latencies, ticket counts, owners), and facts come in three variants:

Compact the way the app does. The conversation is replayed turn by turn through the app's own summarize step, so a long conversation gets summarized several times as it grows. Then each fact is probed with a question such as "Which port does the CDN purge job listen on? Answer with just the value."

Separate the summary from the model. An oracle pseudo-model answers correctly if and only if the value is still in the context it was given. Its recall measures what the summary kept. A real model scoring below it means the information was there but missed. A stricter oracle also requires the value to still be attached to its subject.

Judge in pairs. The same fact is probed under each condition, and conditions are compared with a paired proportion test and a 10-point equivalence margin. I used my own Ordal library for this.

Cap the spend. After an expensive surprise (see below), every run first makes one calibration call per model, dry-runs the whole design for free to count tokens, and refuses to start if the estimate exceeds a hard budget.

Finding 1: a good summary can beat the raw history

With no compaction at all, three small models missed 12–19% of facts in a 23k-token conversation. Most misses were updated facts: gemini-3.5-flash-lite gave the superseded value for 10 of 18 of them.

Compacting with a strong summarizer fixed that. The summary kept every fact and resolved each update to its latest value, so stale answers dropped to zero.

A good summary beat the raw history: all three small models scored higher reading gemini-3.5-flash's summary than the full 23k-token conversation.
A good summary beat the raw history: all three small models scored higher reading gemini-3.5-flash's summary than the full 23k-token conversation.

A summary is not just a shorter transcript. It is a consolidated state, and that can be easier for a model to use than the raw history.

Finding 2: loss compounds with every pass

With a cheap summarizer, the number of summaries a fact has been through predicts whether it survives. After four passes, nothing was left.

Bar chart: facts still in context after 1, 2, 3 and 4 summary passes: 100%, 89%, 65%, 0%.
Facts kept after each summary pass (gemini-3.5-flash-lite summaries; n = 10, 18, 20, 24 facts).

Is it the passes, or a lack of room? Two explanations fit that curve. Either each pass drops a little of what the previous one kept, or a fixed-size summary pushes out the oldest facts. Because older facts have always been through more passes, the two are confounded.

So I summarized the same conversations two ways with the same model: incrementally as they grew, or once at the end.

Summarizing once kept 19 of 24 early facts attached to their subject; summarizing incrementally kept 8.
Summarizing once kept 19 of 24 early facts attached to their subject; summarizing incrementally kept 8.

A single pass kept most of the early facts that incremental summarizing lost. That points to repeated passes rather than lack of room for old facts, though the overall paired test was not decisive (+12.5 points, 95% interval −3.5 to +27.7).

One big pass has its own cost, though. It kept 92% of values somewhere, but only 75% still attached to the right subject. It loses who a value belonged to.

Repetition protects a fact. I restated each fact once, later in the conversation, changing nothing else. Recall rose from 23/72 to 41/72, +25 points (95% interval +12 to +37). Part of that gain is that the restatement is newer, so it went through fewer passes.

Finding 3: the instruction matters as much as the model

The app's original summary prompt asks for "a concise, detailed technical brief" that preserves who argued what. I compared it with Codex CLI's handoff prompt and with a prompt I wrote from first principles, "state and index". That prompt asks for every decision, agreed value and open item with who stated it, only the latest value of anything changed, and a one-line index of topics.

Half the facts in these conversations were stated by the assistant rather than the user.

The instruction matters as much as the model: state-and-index took gemini-3.5-flash-lite from 31 to 70 of 72 facts.
The instruction matters as much as the model: state-and-index took gemini-3.5-flash-lite from 31 to 70 of 72 facts.

Telling a cheap model what kind of thing must survive took it from 31 to 70 facts. It paid for that by writing twice as much. A capable model kept nearly everything even in Codex's roughly 1,300-token handoffs, so the limit is what the summarizer chooses to keep, not the room it has.

What a summary costs. The app's old default summarizer, gemini-3.5-flash, spent 10–20k reasoning tokens per summary at $9 per million output tokens: about $0.21 per summary. gemini-3.8-flash kept the same facts (equivalent within 10 points) for about $0.03, so the app switched to it, until Finding 7 below. Cutting 3.5-flash's thinking to "medium" matched the default; "minimal" cost a quarter but lost 18 points.

A bug surfaced along the way. The app counted the summary itself as new material, so once a summary outgrew the threshold, it re-summarized every few messages: 132 summaries where about 20 were needed. The fix requires half a threshold of genuinely new text before summarizing again.

How Codex CLI and Claude Code do it

Codex CLI is open source, so I read its compaction code. Claude Code isn't, so for it I relied on the official docs.

How Codex CLI and Claude Code compact, next to The Crucible before this study.
How Codex CLI and Claude Code compact, next to The Crucible before this study.

Codex even warns after compacting that "long threads and multiple compactions can cause the model to be less accurate." That is Finding 2, stated by its authors.

I added Codex's rules to the experiment. Keeping user messages verbatim rescued exactly the user-stated facts: with a cheap summarizer, they went from 17/36 to 36/36, while assistant-stated facts did not move. The cost was 2.3 times more context on every later call.

The principle: never let the summary be the only copy

It is tempting to say coding agents and conversational apps simply need different compaction. I think the real distinction is deeper. Every piece of context has three properties:

Read that way, Codex's design makes sense. User messages are authoritative and unrecoverable, so they stay verbatim. Tool output is recoverable from the repository, so it can go. The summary only has to carry live state, which is why a short handoff is enough. Claude Code's CLAUDE.md does the same thing by moving authoritative rules out of the conversation, where compaction can't touch them.

Every loss I measured was information whose only copy was the summary. The Crucible never deletes history; it keeps every message in its checkpoint store. The model just couldn't see it after pruning. So the fix is not a better summary but making the original reachable again.

Finding 4: bringing the original back rescues a lossy summary

I added a recall step. Before each call, it searches the hidden history for the four messages that best match the latest request and quotes them back into context. The search is plain keyword matching weighted by word rarity.

Recall rescued the lossy summaries: 31 → 67 and 9 → 66 of 72 facts.
Recall rescued the lossy summaries: 31 → 67 and 9 → 66 of 72 facts.

Recall added 50 and 79 points to the lossy configurations, and changed nothing where the summary already kept everything. The cheapest setup that kept nearly every detail was the cheapest summarizer writing 1k-token handoffs plus recall: 66/72, with each later call reading 2.8k tokens.

Once the original is reachable, the summary no longer has to be a complete record. It only has to be good enough to continue from.

Finding 5: recall moves the relevance decision to the retriever

Recall only worked so well because my questions named their subject. So I rewrote every question to describe the subject instead: "Which network port should the firewall allow for the task that clears cached files at the edge?" rather than "Which port does the CDN purge job listen on?" No subject word appears in any of them.

With indirect questions, keyword recall fell from 66 to 28 of 72; letting a model choose the search terms recovered 54.
With indirect questions, keyword recall fell from 66 to 28 of 72; letting a model choose the search terms recovered 54.

When the facts are in view, wording doesn't matter: the model maps "the task that clears cached files at the edge" to the CDN purge job almost every time. Recall hands that judgment to the retriever, and a keyword or small-embedding retriever is far worse at it than the model. Letting a cheap model read the summary and choose the search terms recovered most of the loss, but not all.

So both designs fail on capability, just in different places. A summary fails in the summarizer, compounding with every pass. Recall fails in the retriever. Whatever decides relevance has to understand the question about as well as the model does.

A prediction that failed

I expected facts that sound unimportant when they are said to be dropped more often, even by strong summarizers. So I restated every fact as a throwaway remark ("Side note, probably irrelevant: the staging database listens on port 6543"), changing nothing else.

The opposite happened. A cheap summarizer kept more of the asides: 18/72 instead of 1/72 with Codex-style handoffs. The strong summarizer and recall kept everything either way. In filler where every sentence has the same shape, a hedge makes a sentence stand out rather than fade. Sounding unimportant is not the same as being unimportant to a model, and the harder question, relevance nobody could have foreseen when the fact was stated, remains open.

Finding 6: summaries drop what the task doesn't need yet

So I tied relevance to the task instead of the wording. Each conversation pursues one stated goal ("this week's only priority is the log shipper"). Eight facts are about that goal. Eight more are about another subsystem nobody is working on yet, stated plainly in passing, in exactly the same form as everything else, and mentioned nowhere else. The last turn switches the goal: "Change of plans: the warehouse loader is now the priority." Then every fact is probed.

Until this point every summarizer I had tested was a Gemini model, so I ran two from each provider family on four conversations:

After a change of plans, OpenAI and Anthropic summaries kept only 4–13 of the 32 off-goal facts; only gemini-3.8-flash kept most (26/32).
After a change of plans, OpenAI and Anthropic summaries kept only 4–13 of the 32 off-goal facts; only gemini-3.8-flash kept most (26/32).

Four of the four OpenAI and Anthropic summarizers lose off-goal facts, and the OpenAI ones keep almost none. Gemini, the family I had been testing all along, does it least.

My first reading, from two Gemini conversations, was that stronger summarizers drop more. It doesn't hold across families: the stronger model is more selective within OpenAI and less selective within Anthropic. The family's summarizing style matters more than strength. What does hold is the mechanism: a task-focused summary leaves out what the task doesn't need yet, and no summarizer can know the plan is about to change.

Recall brings them all back. Recall looks at the history after the goal has switched, so it should not care about the change of plans. I tested it with the two most goal-selective summarizers:

Recall brought back all 32 off-goal facts for both goal-selective summarizers.
Recall brought back all 32 off-goal facts for both goal-selective summarizers.

Deciding at read time recovers exactly what deciding at write time loses. This is the one result in the study I expect to survive new models: a summary is written for the goal of the moment, no summarizer can anticipate a change of plans, and only a design that keeps the originals reachable is immune. One caveat carries over from Finding 5: these questions named their subject, so keyword recall was enough. With the same questions asked indirectly, keyword recall restored only 16–17 of the 32 and embedding recall did no better, while model-guided recall, where a small model reads what is still visible (including the turn that changed the goal) and chooses what to fetch, restored 28–30. The two summarizers here are two replications, not a comparison: they differ in family and tier.

Choosing what to fetch doesn't need a strong model. Until then the model choosing the search terms had always been gpt-6-luna, the cheapest in the app. So I fixed one summarizer (gpt-6-sol, the most goal-selective), wrote its summaries once, and replayed them with six different models choosing what to fetch, a cheap and a strong one from each family. Same indirect questions, four conversations:

Every model choosing what to fetch beat keyword recall; the cheapest, gpt-6-luna (30/32), matched the strongest.
Every model choosing what to fetch beat keyword recall; the cheapest, gpt-6-luna (30/32), matched the strongest.

Every model beats keyword recall. Luna and sol are statistically equivalent, and the strong models are equivalent to luna; only the smallest Gemini falls somewhat behind. The cheapest model made the call as well as gemini-3.8-flash, which cost 24 times as much for the same calls ($0.59 against $0.024). That fits the time-of-decision story: before the question is known, no summarizer can judge relevance reliably; once it is known, even a small model can. Which cheap model is good enough will change with every release. That deciding late makes the decision easy should not.

Why the summary drops them. "No summarizer can anticipate a change of plans" is a claim, so I tested it. I reran the same conversations with one change: every goal reminder also said "once that is done, the next priority will be X". The facts, filler and positions stayed identical. Separately, I tried a generic hedge that knows nothing about what comes next: the usual instruction plus "priorities may change; keep the specific values for every subject, not only the current goal."

Announcing the next goal took gpt-6-sol from 4 to 30 of 32; the generic hedge worked only with 9× longer summaries.
Announcing the next goal took gpt-6-sol from 4 to 30 of 32; the generic hedge worked only with 9× longer summaries.

For gpt-6-sol, knowing the next goal is the whole story: 4 to 30 of 32, for summaries 30% longer. The generic hedge also kept everything, but only by barely compressing: summaries of 5.5k tokens against an 8k threshold, rewritten twice as often. Without knowing the future, the only way to keep off-goal facts at write time is to give up the compression. Claude Haiku shows a second, separate loss: told the future and hedged, it still kept only 20 of 32, and it lost on-goal facts too. That looks like fidelity, not selection. Read-time recall sidesteps both, because the originals are still there. Was it only that the subject's name kept coming up? Announcing the next goal names it a dozen times, so I ran a control that names it just as often, as out of scope this quarter, with the change of plans still a surprise. gpt-6-sol kept 6 of 32: about 2 of the 26 recovered facts come from the name standing out, and the other 24 from knowing it comes next.

Finding 7: once the originals are reachable, the cheapest summarizer is enough

Recall changes what a summarizer is for, so I chose the app's default summarizer again, this time with recall on, as the app runs. Before running, I wrote down the rule: keep the candidates statistically no worse than the best on both scenarios, then take the cheapest per summary at the prices in force from January 2027 (gemini-3.8-flash's introductory price ends in December). Four summarizers, the detail-heavy scenario and the change-of-plans scenario, four conversations each, strict oracle:

With recall on, gpt-6-luna's summaries matched the best at about 1/25 of the cost; alone they keep much less.
With recall on, gpt-6-luna's summaries matched the best at about 1/25 of the cost; alone they keep much less.

Only gpt-6-luna passed the rule: statistically equivalent to the best on the detail-heavy scenario, itself the best after a change of plans, and about 1/25 of gemini-3.8-flash's cost per summary (costs are for the detail-heavy scenario at 2027 prices). Alone, its summaries keep less than half as much. Once the originals can be fetched, the summary's job shrinks from being the record to being a pointer, and a pointer can be cheap. The choice holds only because recall is on: without it, the ranking inverts.

One result I can't explain yet: recall added 16 to 29 facts for every other summarizer, but nothing for gemini-3.8-flash after the change of plans (56 either way), and a later audit reproduced the gap.

Keeping the scaffolding honest

Every mechanism in this post, from the summary itself to the verbatim window and recall, exists because it beat a baseline on the models I tested. A newer model can close that gap on its own, and scaffolding that no longer helps can get in its way. So each mechanism on the app's default path now declares its baseline and its claim, and an audit re-measures the pair on current models whenever the model list changes, recording keep, retire, harmful or inconclusive.

The first two audits already changed the picture. Both used two cheap answering models on the detail-heavy scenario:

Recall went from a mixed blessing (first audit, thorough summaries) to essential (second audit, cheap summaries).
Recall went from a mixed blessing (first audit, thorough summaries) to essential (second audit, cheap summaries).

With a thorough summary, the recalled excerpts were a distraction for the weaker model: it took values from the wrong subject and gave the same answer to different questions. With a thin summary, the excerpts are the main source, and recall became essential for both. The second audit also checked the summarizer choice with real answering models: neither pricier summarizer did better for either reader.

Not every verdict should be acted on. The audit flagged the verbatim window, which keeps the newest messages as they are, as "retire": with recall on, it added nothing to fact recall. But the probes only ask about planted facts, and the window exists for continuity with the last few turns, which they don't measure. An audit is only as good as its claim.

What changed in the app

What it cost, and what went wrong

The whole study cost roughly $80 in API calls. About $41 of that went on the first four runs, before I had a budget guard. The rest, about twenty more experiments, cost about $36 under hard caps, plus about $3 that a cost estimate spent by mistake, including one run that returned nothing and a few dollars of calibration for runs that were then refused.

The mistakes were as instructive as the results:

A capped run that returned nothing. A long-horizon run (about 22 summaries per conversation) spent its whole $5 cap and produced zero results. Every condition was replaying at once, the summaries grew with every fold faster than my one-call calibration predicted, and the cap stopped all of them mid-way. Runs now go one condition at a time, cheapest first, each admitted only if the remaining budget covers it and saved as soon as it finishes.

An estimate that spent money. The cost estimate is a dry run with stand-in models, so it should be free. But the model choosing what to fetch was handed the real model client when the experiment was set up, and the dry run never swapped it out. Estimating the six-model run above made about 384 real calls, roughly $3, recorded nowhere and priced at zero. Now each step that calls a model says which model it calls, and the estimate meters it with a stand-in.

The cache that never hit. Cache hits across the experiments were about 1%. The app's system prompt began with the current time to the second, so no two requests shared a prefix. Moving the time out of the system prompt was not enough, because each provider caches differently:

Where per-request context has to go for prompt caching to work, by provider.
Where per-request context has to go for prompt caching to work, by provider.

The general rule: anything that changes per request (a timestamp, retrieved excerpts) belongs after the last point the next request will repeat, and where that point is depends on the provider.

Related work

These results sit next to a growing body of work, and several papers point the same way:

What I haven't seen elsewhere, though I may have missed it: a controlled once-versus-incremental comparison to separate per-pass loss from capacity, an oracle that separates "the summary dropped it" from "the model missed it", a paired test of repetition, a causal test of why task-focused summaries drop facts (announcing the next goal versus only naming it), and an audit that re-checks each piece of context machinery against going without it whenever the models change.

Limitations and what's next

This is a pilot study, not a benchmark:

Next I want to run it on real conversation histories, and add stronger models. The goal is to turn this into a small open-source context-management layer that decides what to keep verbatim, what to summarize and what to leave retrievable.

Code and data. Everything runs in The Crucible (crucible compaction-eval and crucible scaffold-audit). Every number here is in the lab notes, with the raw results next to them.

How this was made. I ran this study with Claude Code as a coding and analysis assistant. It wrote much of the experiment code and the first draft of this post. I directed the questions and checked the results, and I'm responsible for the conclusions.

壓縮之後,還剩下什麼?

一段很長的 LLM 對話為了塞進 context window 而被摘要時,哪些細節會留下來?其他的又為什麼消失?我在很長的合成對話裡埋進事實,用不同方式壓縮,再測量模型還能召回多少。簡單說:摘要會丟掉它無從得知日後會重要的東西,解法是讓原文一直拿得到。

重點摘要

問題

每個長時間運作的 LLM 應用,遲早都會碰到 context window 的上限。常見的解法是壓縮:把較早的歷史摘要起來,接著從摘要繼續。Claude Code 這樣做,Codex CLI 也這樣做,我維護的開源多模型聊天應用 The Crucible 也是。

我的問題很簡單:模型會不會在某個點突然忘記?如果細節真的遺失了,原因是什麼?是摘要、模型的注意力,還是反覆壓縮這件事本身?

我找不到有對照的答案,所以自己在應用裡做了一套。以下所有實驗都跑在這個應用真正的 LangGraph 流程上,程式碼和原始結果都已公開。

我怎麼測量

埋進去的事實。每段合成對話有 140 輪往返,大約 23k tokens,其中 24 個虛構的事實藏在訊息中段:預算上限、審查者、金鑰標籤、埠號、日期、客戶原話的引述、供應商選擇,以及「我們說好不用 X」。因為數值是虛構的,模型無法靠既有知識作答。

讓細節互相競爭。我的第一次小規模試驗在所有條件下都拿到 100%,因為在一般性的填充內容裡,這些事實是唯一具體的內容。所以現在每一輪都帶有形式相同的具體資訊(延遲、工單數、負責人),而事實分成三種變體:

照應用的方式壓縮。對話會一輪一輪地重播,經過應用本身的摘要步驟,所以長對話在變長的過程中會被摘要好幾次。接著用問題逐一探測每個事實,例如「CDN purge job 監聽哪個埠?只回答數值。」

把摘要和模型分開。我用一個 oracle(只看資訊在不在 context 裡的假模型):若且唯若數值還在給它的 context 裡,它就答對。它的召回率反映的是摘要保留了什麼。真實模型的分數如果低於它,代表資訊明明在,卻被漏掉了。更嚴格的 oracle 還要求數值仍然跟它所屬的對象連在一起。

成對比較。每個事實在每種條件下都探測一次,條件之間用成對比例檢定和 10 個百分點的等價界限來比較。這部分我用的是自己寫的 Ordal 函式庫。

限制花費。在一次昂貴的意外之後(見下文),每次執行都會先對每個模型做一次校準呼叫,免費空跑整個實驗設計來計算 token 數,如果估計值超過硬性預算就拒絕開始。

發現 1:好的摘要可能勝過原始歷史

完全不壓縮時,在一段 23k tokens 的對話裡,三個小模型漏掉了 12–19% 的事實。漏掉的大多是更新過的事實:18 個更新過的事實中,gemini-3.5-flash-lite 有 10 個給出了已被取代的舊值。

用強的摘要模型壓縮就解決了這個問題。摘要保留了所有事實,並把每次更新都整理成最新的值,過時的答案因此降到零。

好的摘要勝過原始歷史:三個小模型讀 gemini-3.5-flash 的摘要時,得分都高於讀完整的 23k tokens 對話。
好的摘要勝過原始歷史:三個小模型讀 gemini-3.5-flash 的摘要時,得分都高於讀完整的 23k tokens 對話。

摘要不只是比較短的逐字稿。它是整合過的狀態,而對模型來說,這可能比原始歷史更好用。

發現 2:每多摘要一次,損失就多累積一次

用便宜的摘要模型時,一個事實經過幾次摘要,就能預測它會不會留下來。經過四次之後,什麼都不剩。

長條圖:經過 1、2、3、4 次摘要後仍留在 context 中的事實:100%、89%、65%、0%。
每次摘要後保留下來的事實(gemini-3.5-flash-lite 的摘要;n = 10、18、20、24 個事實)。

是摘要次數的問題,還是空間不夠?有兩種解釋都符合這條曲線。一是每次摘要都會丟掉一點上一次保留的東西;二是固定大小的摘要會把最舊的事實擠出去。因為較舊的事實一定經過比較多次摘要,這兩個因素彼此混淆,分不開。

所以我用同一個模型,以兩種方式摘要同樣的對話:隨著對話變長逐步摘要,或是在最後一次摘要完。

一次摘要完,24 個早期事實中有 19 個仍跟所屬對象連在一起;逐步摘要只剩 8 個。
一次摘要完,24 個早期事實中有 19 個仍跟所屬對象連在一起;逐步摘要只剩 8 個。

單次摘要保留了大部分逐步摘要弄丟的早期事實。這指向的是反覆摘要,而不是舊事實沒有空間;不過整體的成對檢定並不是決定性的(+12.5 個百分點,95% 信賴區間 −3.5 到 +27.7)。

不過,一次大摘要也有自己的代價。它把 92% 的數值保留在某個地方,但只有 75% 還跟正確的對象連在一起。它弄丟的是一個數值屬於誰。

重複能保護一個事實。我在對話稍後把每個事實再講一次,其他什麼都沒改。召回從 23/72 上升到 41/72,增加 25 個百分點(95% 信賴區間 +12 到 +37)。這項提升有一部分是因為重述比較新,所以經過的摘要次數比較少。

發現 3:指示跟模型一樣重要

應用原本的摘要 prompt 要求「一份精簡而詳盡的技術簡報」,並保留誰主張了什麼。我拿它跟 Codex CLI 的交接 prompt,以及一個我從基本原理出發寫的「狀態與索引」(state and index)prompt 比較。這個 prompt 要求列出每個決定、議定的數值和待辦事項,並註明是誰提出的;任何被改過的東西只留最新的值;再加上一行主題索引。

這些對話裡,有一半的事實是助理說的,而不是使用者。

指示跟模型一樣重要:狀態與索引讓 gemini-3.5-flash-lite 從 72 個事實中的 31 個提升到 70 個。
指示跟模型一樣重要:狀態與索引讓 gemini-3.5-flash-lite 從 72 個事實中的 31 個提升到 70 個。

告訴便宜模型必須留下的是哪一類東西,就讓它從 31 個事實提升到 70 個。代價是它寫出來的長度多了一倍。能力夠的模型即使在 Codex 大約 1,300 tokens 的交接摘要裡,也幾乎全部保留,所以限制在於摘要模型選擇保留什麼,而不是它有多少空間。

一份摘要要花多少錢。應用原本的預設摘要模型 gemini-3.5-flash,每份摘要會花 10–20k 個推理 token,輸出 token 每百萬 $9:每份摘要大約 $0.21。gemini-3.8-flash 保留了同樣的事實(在 10 個百分點內等價),只要大約 $0.03,所以應用改用它,直到下文的發現 7 為止。把 3.5-flash 的思考程度降到「medium」,表現跟預設設定相當;「minimal」的成本只有四分之一,但少了 18 個百分點。

過程中還冒出一個 bug。應用把摘要本身算成新內容,所以摘要一旦長到超過門檻,每隔幾則訊息就會重新摘要一次:本來大約 20 份就夠,結果產生了 132 份。修正方式是要求累積到半個門檻量的真正新文字,才再次摘要。

Codex CLI 和 Claude Code 怎麼做

Codex CLI 是開源的,所以我讀了它的壓縮程式碼。Claude Code 不是開源的,所以我依據的是它的官方文件。

Codex CLI 和 Claude Code 的壓縮方式,與這項研究之前的 The Crucible 並列比較。
Codex CLI 和 Claude Code 的壓縮方式,與這項研究之前的 The Crucible 並列比較。

Codex 甚至會在壓縮後警告:「長對話和多次壓縮可能讓模型的準確度下降」(long threads and multiple compactions can cause the model to be less accurate)。這就是發現 2,只是由它的作者親口說出來。

我把 Codex 的規則加進實驗。原文保留使用者訊息,救回的恰好就是使用者說出的事實:用便宜的摘要模型時,這類事實從 17/36 提升到 36/36,而助理說出的事實完全沒變。代價是之後每次呼叫都要多用 2.3 倍的 context。

原則:永遠別讓摘要成為唯一的副本

很容易就會下結論說,程式碼代理和對話型應用本來就需要不同的壓縮方式。我認為真正的差別更深一層。每一段 context 都有三個屬性:

這樣看,Codex 的設計就說得通了。使用者訊息具有權威性又無法復原,所以原文保留。工具輸出可以從 repository 重新取得,所以可以丟掉。摘要只需要承載仍然有效的狀態,這就是為什麼簡短的交接摘要就夠了。Claude Code 的 CLAUDE.md 做的是同一件事:把權威性的規則移出對話,放到壓縮碰不到的地方。

我測到的每一次損失,都是唯一副本只存在於摘要裡的資訊。The Crucible 從不刪除歷史;每則訊息都保存在它的 checkpoint store 裡。只是修剪之後,模型看不到它們而已。所以解法不是更好的摘要,而是讓原文重新拿得到。

發現 4:把原文找回來,能救回有損的摘要

我加了一個召回步驟。每次呼叫之前,它會在隱藏的歷史中搜尋跟最新請求最相符的四則訊息,把它們引用回 context 裡。搜尋方式是單純的關鍵字比對,依字詞的稀有程度加權。

召回救回了有損的摘要:72 個事實中,31 → 67,以及 9 → 66。
召回救回了有損的摘要:72 個事實中,31 → 67,以及 9 → 66。

召回讓兩個有損的配置分別提升了 50 和 79 個百分點,而在摘要本來就保留了一切的配置上沒有任何改變。能保留幾乎所有細節的最便宜設定,是最便宜的摘要模型寫 1k tokens 的交接摘要,再加上召回:66/72,之後每次呼叫讀 2.8k tokens。

一旦原文拿得到,摘要就不必是完整的紀錄,只要好到足以接續下去就行。

發現 5:召回把「什麼才相關」的判斷交給了檢索器

召回之所以效果這麼好,只是因為我的問題直接點名了主題。所以我把每個問題都改寫成描述主題,而不直接說出來:問「防火牆應該為那個在邊緣清除快取檔案的任務開放哪個網路埠?」,而不是「CDN purge job 監聽哪個埠?」。這些問題裡都沒有出現任何主題詞。

改用間接問題後,關鍵字召回從 72 個中的 66 個掉到 28 個;讓模型選擇搜尋詞,則救回到 54 個。
改用間接問題後,關鍵字召回從 72 個中的 66 個掉到 28 個;讓模型選擇搜尋詞,則救回到 54 個。

事實就在眼前時,措辭不重要:模型幾乎每次都能把「在邊緣清除快取檔案的任務」對應到 CDN purge job。召回把這個判斷交給檢索器,而關鍵字或小型 embedding 檢索器在這方面遠不如模型。讓便宜的模型讀摘要、選擇搜尋詞,能救回大部分的損失,但不是全部。

所以兩種設計都會因為能力不足而失敗,只是失敗在不同的地方。摘要失敗在摘要模型,而且每多摘要一次就多累積一次。召回失敗在檢索器。不管由誰來決定相關性,它對問題的理解都得跟模型差不多好。

一個落空的預測

我原本預期,說出口時聽起來不重要的事實,會更常被丟掉,即使是強的摘要模型也一樣。所以我把每個事實都改寫成隨口一提的話(「順帶一提,可能無關:staging 資料庫監聽 6543 埠」),其他什麼都沒改。

結果正好相反。便宜的摘要模型反而保留了更多這種順帶一提的話:用 Codex 式的交接摘要時是 18/72,而不是 1/72。強的摘要模型和召回則不論哪種寫法都全部保留。在每句話形式都一樣的填充內容裡,帶保留語氣的句子反而更突出,而不是被淹沒。聽起來不重要,不等於對模型來說不重要;而更難的問題,也就是事實被說出時沒有人能預見的相關性,仍然沒有答案。

發現 6:摘要會丟掉任務暫時還不需要的東西

於是我改把相關性綁在任務上,而不是措辭上。每段對話都追求一個明確的目標(「這週唯一的優先事項是 log shipper」)。有八個事實跟這個目標有關。另外八個是關於另一個目前還沒有人在處理的子系統,只是順口平鋪直敘地帶過,形式跟其他內容完全相同,而且其他地方都沒有再提到。最後一輪切換目標:「改變計畫:現在優先處理 warehouse loader。」接著探測每一個事實。

在這之前,我測過的摘要模型全都是 Gemini,所以我從每個供應商家族各挑兩個,在四段對話上測試:

改變計畫之後,OpenAI 和 Anthropic 的摘要只保留了 32 個目標外事實中的 4–13 個;只有 gemini-3.8-flash 保留了大部分(26/32)。
改變計畫之後,OpenAI 和 Anthropic 的摘要只保留了 32 個目標外事實中的 4–13 個;只有 gemini-3.8-flash 保留了大部分(26/32)。

OpenAI 和 Anthropic 的四個摘要模型,四個全都會丟掉目標外的事實,而 OpenAI 的兩個幾乎一個都沒留。我一直以來在測的 Gemini 家族,丟得最少。

我最初根據兩段 Gemini 對話得出的解讀是:越強的摘要模型丟得越多。但這在不同家族之間並不成立:在 OpenAI 家族裡,較強的模型篩選得比較嚴;在 Anthropic 家族裡,較強的模型反而篩選得比較鬆。家族的摘要風格比模型強弱更重要。真正站得住的是機制:以任務為中心的摘要,會省略任務暫時還不需要的東西,而沒有任何摘要模型能知道計畫即將改變。

召回把它們全部找回來。召回是在目標切換之後才去看歷史,所以照理說不該受到改變計畫的影響。我用兩個最會依目標篩選的摘要模型來測試:

對兩個依目標篩選的摘要模型,召回都找回了全部 32 個目標外事實。
對兩個依目標篩選的摘要模型,召回都找回了全部 32 個目標外事實。

在讀取時才決定,恰好能找回在寫入時就決定所損失的東西。這是整個研究中,我唯一預期在新模型出來後仍會成立的結果:摘要是為當下的目標而寫,沒有摘要模型能預料到計畫改變,只有讓原文保持可取得的設計不受影響。有一個但書延續自發現 5:這些問題都點名了主題,所以關鍵字召回就夠了。同樣的問題改成間接問法時,關鍵字召回只找回 32 個中的 16–17 個,embedding 召回也沒有比較好;而模型引導的召回(由小模型讀取仍然看得到的內容,包括改變目標的那一輪,再選擇要取回什麼)則找回了 28–30 個。這裡的兩個摘要模型是兩次重複驗證,不是一組比較:它們的家族和等級都不同。

選擇要取回什麼,不需要強模型。在那之前,負責選搜尋詞的模型一直是應用裡最便宜的 gpt-6-luna。所以我固定一個摘要模型(最會依目標篩選的 gpt-6-sol),只寫一次摘要,再重播給六個不同的模型,讓它們選擇要取回什麼,每個家族各一個便宜的、一個強的。同樣的間接問題,四段對話:

每個負責選擇取回內容的模型都勝過關鍵字召回;最便宜的 gpt-6-luna(30/32)追平了最強的模型。
每個負責選擇取回內容的模型都勝過關鍵字召回;最便宜的 gpt-6-luna(30/32)追平了最強的模型。

每個模型都勝過關鍵字召回。Luna 和 sol 在統計上等價,強模型也都跟 luna 等價;只有最小的 Gemini 稍微落後。最便宜的模型判斷得跟 gemini-3.8-flash 一樣好,而後者做同樣的呼叫要花 24 倍的錢($0.59 對 $0.024)。這符合「決定的時機」這個解釋:問題確定之前,沒有摘要模型能可靠地判斷相關性;問題一旦確定,連小模型都做得到。哪個便宜模型夠用,每次新版本推出都會變;但「晚點決定會讓決定變簡單」這件事應該不會變。

摘要為什麼會丟掉它們。「沒有摘要模型能預料到計畫改變」是一個主張,所以我測試了它。我重跑同樣的對話,只改一個地方:每次提醒目標時,都多加一句「完成之後,下一個優先事項會是 X」。事實、填充內容和位置都完全一樣。另外,我也試了一個對接下來會發生什麼一無所知的通用保險措施:在平常的指示之外再加上「優先順序可能改變;每個主題的具體數值都要保留,不只是目前的目標。」

宣告下一個目標,讓 gpt-6-sol 從 32 個中的 4 個提升到 30 個;通用保險措施只有在摘要長了 9× 時才有效。
宣告下一個目標,讓 gpt-6-sol 從 32 個中的 4 個提升到 30 個;通用保險措施只有在摘要長了 9× 時才有效。

對 gpt-6-sol 來說,知道下一個目標就是全部的關鍵:32 個中從 4 個提升到 30 個,代價是摘要長了 30%。通用保險措施也保留了所有事實,但靠的是幾乎不壓縮:摘要有 5.5k tokens,而門檻是 8k,重寫的頻率也多了一倍。不知道未來的話,在寫入時保留目標外事實的唯一方法,就是放棄壓縮。Claude Haiku 則顯示出第二種、另一回事的損失:即使告訴它未來、也加了保險措施,它仍然只保留 32 個中的 20 個,而且連目標內的事實也丟了。這看起來是保真度的問題,而不是篩選的問題。讀取時的召回兩者都能避開,因為原文還在。會不會只是因為那個主題的名稱一直被提到?宣告下一個目標會提到它十幾次,所以我做了一個對照組:提到它的次數一樣多,但說它是這一季不處理的範圍,改變計畫則仍然出乎意料。gpt-6-sol 保留了 32 個中的 6 個:救回的 26 個事實裡,大約 2 個來自名稱變得醒目,其餘 24 個來自知道它是下一個目標。

發現 7:原文拿得到時,最便宜的摘要模型就夠用

召回改變了摘要模型的用途,所以我重新挑選應用的預設摘要模型,這次開啟召回,跟應用實際運作時一樣。執行之前,我先寫下規則:留下在兩種情境下在統計上都不比最好的差的候選,再依 2027 年 1 月起適用的價格,選出每份摘要最便宜的那個(gemini-3.8-flash 的優惠價到 12 月結束)。四個摘要模型、細節密集的情境和改變計畫的情境,各四段對話,用嚴格的 oracle:

開啟召回時,gpt-6-luna 的摘要以大約 1/25 的成本追平了最好的;單獨使用時保留的內容少得多。
開啟召回時,gpt-6-luna 的摘要以大約 1/25 的成本追平了最好的;單獨使用時保留的內容少得多。

只有 gpt-6-luna 通過了這條規則:在細節密集的情境上跟最好的在統計上等價,在改變計畫之後本身就是最好的,而且每份摘要的成本大約只有 gemini-3.8-flash 的 1/25(成本以細節密集的情境、2027 年的價格計算)。單獨使用時,它的摘要保留的內容不到一半。一旦原文可以取回,摘要的工作就從「作為紀錄」縮小成「作為指標」,而指標可以很便宜。這個選擇只在召回開啟時成立:沒有召回,排名就會反過來。

有一個結果我目前還無法解釋:召回讓其他每個摘要模型都多保住了 16 到 29 個事實,但在改變計畫之後,對 gemini-3.8-flash 卻毫無幫助(兩種情況都是 56),而之後的一次稽核也重現了這個落差。

讓輔助機制保持誠實

這篇文章裡的每一個機制,從摘要本身到原文保留視窗和召回,之所以存在,都是因為它在我測試的模型上勝過了基準線。較新的模型可能自己就把這個差距補上了,而不再有幫助的輔助機制可能反而礙事。所以應用預設路徑上的每個機制,現在都會宣告自己的基準線和主張;每當模型清單改變,稽核就會在目前的模型上重新測量這一組比較,並記錄為保留、淘汰、有害或無定論。

前兩次稽核就已經改變了整個局面。兩次都用兩個便宜的作答模型,跑細節密集的情境:

召回從好壞參半(第一次稽核,周全的摘要)變成不可或缺(第二次稽核,便宜的摘要)。
召回從好壞參半(第一次稽核,周全的摘要)變成不可或缺(第二次稽核,便宜的摘要)。

摘要很周全時,召回的片段對較弱的模型反而是干擾:它會拿錯對象的數值,對不同的問題給出同一個答案。摘要很單薄時,這些片段就是主要來源,召回對兩個模型都變得不可或缺。第二次稽核也用真實的作答模型檢查了摘要模型的選擇:兩個較貴的摘要模型,對哪一個讀者都沒有表現得更好。

不是每個判定都該照做。稽核把原文保留視窗(原樣保留最新幾則訊息)標為「淘汰」:開啟召回後,它對事實召回毫無幫助。但這些探測只問埋進去的事實,而這個視窗存在的目的是跟最近幾輪保持連貫,這一點探測並沒有測到。稽核的好壞,取決於它檢驗的主張。

應用裡改了什麼

花了多少錢,又出了哪些錯

整個研究的 API 呼叫大約花了 $80。其中約 $41 花在最初的四次執行,那時我還沒有預算防護。其餘大約二十個實驗在硬性上限下花了約 $36,其中包括一次什麼結果都沒產出的執行,以及幾塊錢花在後來被拒絕執行的實驗的校準上;另外還有約 $3 是成本估算誤花掉的。

犯的錯跟結果一樣有啟發性:

一次設了上限卻什麼都沒產出的執行。一次長程執行(每段對話大約 22 份摘要)花光了整整 $5 的上限,卻產出零筆結果。所有條件同時在重播,摘要每合併一次就變長,變長的速度比我用單次呼叫校準的預測還快,結果上限讓所有條件都在中途停下。現在一次只跑一個條件,從最便宜的開始,只有剩餘預算足夠時才放行,而且每個條件一跑完就立刻存檔。

一次會花錢的估算。成本估算是用替身模型空跑,所以應該是免費的。但負責選擇取回內容的模型,在設定實驗時就被交給了真正的模型 client,而空跑從來沒有把它換掉。估算上面那個六模型的實驗時,發出了大約 384 次真實呼叫,約 $3,哪裡都沒有紀錄,還被計價為零。現在每個會呼叫模型的步驟都會說明它呼叫哪個模型,估算時則用替身來計量。

從沒命中的快取。整個實驗期間的快取命中率大約只有 1%。應用的 system prompt 一開頭就是精確到秒的當下時間,所以任兩個請求都沒有共同的前綴。把時間移出 system prompt 還不夠,因為每家供應商的快取方式都不一樣:

依供應商整理:每次請求才有的 context 要放在哪裡,prompt caching 才會生效。
依供應商整理:每次請求才有的 context 要放在哪裡,prompt caching 才會生效。

通則是:任何每次請求都會變的東西(時間戳記、檢索到的片段),都應該放在下一個請求會重複的最後一個位置之後,而那個位置在哪裡,取決於供應商。

相關研究

這些結果跟一批越來越多的研究放在一起看,其中好幾篇指向同一個方向:

我在其他地方還沒看到的(不過也可能是我漏看了):一組有對照的「一次摘要」與「逐步摘要」比較,用來區分每次摘要的損失和容量不足;一個能把「摘要丟掉了」跟「模型漏看了」分開的 oracle;一個針對重複的成對檢定;一個探究以任務為中心的摘要為什麼會丟掉事實的因果檢驗(宣告下一個目標,對比只提到它的名稱);以及一套稽核,每當模型改變,就把每一個處理 context 的機制拿來跟「不用它」重新比較。

限制與下一步

這是一項小規模試驗,不是基準測試:

接下來我想在真實的對話歷史上跑這套實驗,並加入更強的模型。目標是把它變成一個小型的開源 context 管理層,決定哪些要原文保留、哪些要摘要、哪些留著可供檢索。

程式碼與資料。所有東西都在 The Crucible 裡執行(crucible compaction-eval 和 crucible scaffold-audit)。這裡的每個數字都記錄在實驗筆記中,原始結果就放在旁邊。

這篇是怎麼做出來的。我在這項研究中用 Claude Code 當作寫程式和分析的助手。很多實驗程式碼,以及這篇文章的初稿,都是它寫的。我負責提出問題、檢查結果,結論由我負責。

Knas 的文章。想聊聊的話:Written by Knas. To get in touch: Email GitHub LinkedIn