<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom">
<channel>
  <title>Writing by Knas</title>
  <link>https://www.imaknas.com/#writing</link>
  <atom:link href="https://www.imaknas.com/writing/feed.en.xml" rel="self" type="application/rss+xml"/>
  <description>Writing by Knas</description>
  <language>en</language>
  <item>
    <title>What Survives Compaction?</title>
    <link>https://www.imaknas.com/writing/what-survives-compaction/?lang=en</link>
    <guid isPermaLink="false">https://www.imaknas.com/writing/what-survives-compaction/#en</guid>
    <pubDate>Mon, 28 Sep 2026 12:00:00 +0000</pubDate>
    <description>When a long LLM conversation is summarized to fit the context window, which details survive and why the rest disappear. A pilot study run in The Crucible.</description>
    <content:encoded><![CDATA[<p><em>When a long LLM conversation gets summarized to fit the context window, which details survive, and why do the rest disappear? I planted facts in long synthetic conversations, compacted them in different ways, and measured what the model could still recall.</em><em> The short version: a summary loses what it can&#x27;t know will matter, and the fix is to keep the originals within reach.</em></p>
<h3>TL;DR</h3>
<ul><li><strong>A summary can&#x27;t know what the next question will be.</strong> After a change of plans, OpenAI and Anthropic summaries had kept only 4–13 of the 32 facts the new goal needed. Announcing the next goal in advance restored 30 of 32; naming it just as often without making it next restored 6. Without that knowledge, the only fix at write time is to stop compressing.</li><li><strong>So never let the summary be the only copy.</strong> Keeping the original messages and letting a small model pick what to fetch once the question is known brought back 28–32 of the 32. The cheapest model chose as well as the strongest.</li><li><strong>With recall on, the cheapest summarizer is enough.</strong> Alone, its summaries keep less than half of what the most thorough one&#x27;s do; with recall they match it, at about 1/25 of the cost per summary.</li><li><strong>Loss happens in the summary, and it compounds.</strong> When the summary dropped a fact, the model couldn&#x27;t answer. With a cheap summarizer, facts that went through one summary always survived; after four, none did.</li><li><strong>What you ask the summarizer to keep matters as much as which model you use.</strong> One instruction change took a cheap model from 31/72 facts to 70/72. And a good summary can beat the raw history: it resolves superseded values that small models otherwise pick.</li><li><strong>Scaffolding has to keep earning its place.</strong> Every mechanism between the conversation and the model is re-measured against going without it. The first audit caught recall misleading a weak model; after the summarizer changed, the second found recall essential.</li><li><strong>Watch the bill.</strong> Reasoning tokens were the biggest cost, my early estimates were off by 2x, and a timestamp at the top of the prompt had silently disabled prompt caching.</li></ul>
<h3>The question</h3>
<p>Every long-running LLM app eventually hits the context window. The common fix is compaction: summarize the older history, keep going from the summary. Claude Code does it, Codex CLI does it, and so does The Crucible, an open-source multi-model chat app I maintain.</p>
<p>My question was simple: is there a point where the model suddenly forgets? And if details do get lost, why? Is it the summary, the model&#x27;s attention, or something about doing it repeatedly?</p>
<p>I couldn&#x27;t find a controlled answer, so I built one into the app. Everything below runs on the app&#x27;s real LangGraph pipeline, and the code and raw results are public.</p>
<h3>How I measured it</h3>
<p><strong>Planted facts.</strong> Each synthetic conversation is 140 exchanges, about 23k tokens, with 24 fictional facts buried mid-message: budget caps, reviewers, key labels, ports, dates, quotes, vendor choices, and &quot;we agreed not to use X&quot;. Fictional values mean the model can&#x27;t answer from prior knowledge.</p>
<p><strong>Make details compete.</strong> My first pilot scored 100% everywhere because the facts were the only specific content in generic filler. So every turn now carries specifics of the same shape (latencies, ticket counts, owners), and facts come in three variants:</p>
<ul><li><strong>Plain:</strong> stated once.</li><li><strong>Distractor:</strong> a rejected proposal with a different value sits nearby.</li><li><strong>Update:</strong> an older value appears first and is later changed. Answering with the old or rejected value counts as wrong.</li></ul>
<p><strong>Compact the way the app does.</strong> The conversation is replayed turn by turn through the app&#x27;s own summarize step, so a long conversation gets summarized several times as it grows. Then each fact is probed with a question such as &quot;Which port does the CDN purge job listen on? Answer with just the value.&quot;</p>
<p><strong>Separate the summary from the model.</strong> An oracle pseudo-model answers correctly if and only if the value is still in the context it was given. Its recall measures what the summary kept. A real model scoring below it means the information was there but missed. A stricter oracle also requires the value to still be attached to its subject.</p>
<p><strong>Judge in pairs.</strong> The same fact is probed under each condition, and conditions are compared with a paired proportion test and a 10-point equivalence margin. I used my own <a href="https://github.com/imaknas/ordal">Ordal</a> library for this.</p>
<p><strong>Cap the spend.</strong> After an expensive surprise (see below), every run first makes one calibration call per model, dry-runs the whole design for free to count tokens, and refuses to start if the estimate exceeds a hard budget.</p>
<h3>Finding 1: a good summary can beat the raw history</h3>
<p>With no compaction at all, three small models missed 12–19% of facts in a 23k-token conversation. Most misses were updated facts: gemini-3.5-flash-lite gave the superseded value for 10 of 18 of them.</p>
<p>Compacting with a strong summarizer fixed that. The summary kept every fact and resolved each update to its latest value, so stale answers dropped to zero.</p>
<figure><img src="https://www.imaknas.com/writing/what-survives-compaction/images/table-01.png" width="600" height="168" class="keep-legible" alt="A good summary beat the raw history: all three small models scored higher reading gemini-3.5-flash's summary than the full 23k-token conversation."><figcaption>A good summary beat the raw history: all three small models scored higher reading gemini-3.5-flash's summary than the full 23k-token conversation.</figcaption></figure>
<p>A summary is not just a shorter transcript. It is a consolidated state, and that can be easier for a model to use than the raw history.</p>
<h3>Finding 2: loss compounds with every pass</h3>
<p>With a cheap summarizer, the number of summaries a fact has been through predicts whether it survives. After four passes, nothing was left.</p>
<figure><img src="https://www.imaknas.com/writing/what-survives-compaction/images/chart-decay.png" width="660" height="380" class="keep-legible" alt="Bar chart: facts still in context after 1, 2, 3 and 4 summary passes: 100%, 89%, 65%, 0%."><figcaption>Facts kept after each summary pass (gemini-3.5-flash-lite summaries; n = 10, 18, 20, 24 facts).</figcaption></figure>
<p><strong>Is it the passes, or a lack of room?</strong> Two explanations fit that curve. Either each pass drops a little of what the previous one kept, or a fixed-size summary pushes out the oldest facts. Because older facts have always been through more passes, the two are confounded.</p>
<p>So I summarized the same conversations two ways with the same model: incrementally as they grew, or once at the end.</p>
<figure><img src="https://www.imaknas.com/writing/what-survives-compaction/images/table-02.png" width="548" height="133" class="keep-legible" alt="Summarizing once kept 19 of 24 early facts attached to their subject; summarizing incrementally kept 8."><figcaption>Summarizing once kept 19 of 24 early facts attached to their subject; summarizing incrementally kept 8.</figcaption></figure>
<p>A single pass kept most of the early facts that incremental summarizing lost. That points to repeated passes rather than lack of room for old facts, though the overall paired test was not decisive (+12.5 points, 95% interval −3.5 to +27.7).</p>
<p>One big pass has its own cost, though. It kept 92% of values somewhere, but only 75% still attached to the right subject. It loses who a value belonged to.</p>
<p><strong>Repetition protects a fact.</strong> I restated each fact once, later in the conversation, changing nothing else. Recall rose from 23/72 to 41/72, +25 points (95% interval +12 to +37). Part of that gain is that the restatement is newer, so it went through fewer passes.</p>
<h3>Finding 3: the instruction matters as much as the model</h3>
<p>The app&#x27;s original summary prompt asks for &quot;a concise, detailed technical brief&quot; that preserves who argued what. I compared it with Codex CLI&#x27;s handoff prompt and with a prompt I wrote from first principles, &quot;state and index&quot;. That prompt asks for every decision, agreed value and open item with who stated it, only the latest value of anything changed, and a one-line index of topics.</p>
<p>Half the facts in these conversations were stated by the assistant rather than the user.</p>
<figure><img src="https://www.imaknas.com/writing/what-survives-compaction/images/table-03.png" width="699" height="273" class="keep-legible" alt="The instruction matters as much as the model: state-and-index took gemini-3.5-flash-lite from 31 to 70 of 72 facts."><figcaption>The instruction matters as much as the model: state-and-index took gemini-3.5-flash-lite from 31 to 70 of 72 facts.</figcaption></figure>
<p>Telling a cheap model <em>what kind</em> of thing must survive took it from 31 to 70 facts. It paid for that by writing twice as much. A capable model kept nearly everything even in Codex&#x27;s roughly 1,300-token handoffs, so the limit is what the summarizer chooses to keep, not the room it has.</p>
<p><strong>What a summary costs.</strong> The app&#x27;s old default summarizer, gemini-3.5-flash, spent 10–20k reasoning tokens per summary at $9 per million output tokens: about $0.21 per summary. gemini-3.8-flash kept the same facts (equivalent within 10 points) for about $0.03, so the app switched to it, until Finding 7 below. Cutting 3.5-flash&#x27;s thinking to &quot;medium&quot; matched the default; &quot;minimal&quot; cost a quarter but lost 18 points.</p>
<p>A bug surfaced along the way. The app counted the summary itself as new material, so once a summary outgrew the threshold, it re-summarized every few messages: 132 summaries where about 20 were needed. The fix requires half a threshold of genuinely new text before summarizing again.</p>
<h3>How Codex CLI and Claude Code do it</h3>
<p>Codex CLI is open source, so I read its <a href="https://github.com/openai/codex/blob/9db8162d65/codex-rs/core/src/compact.rs">compaction code</a>. Claude Code isn&#x27;t, so for it I relied on the <a href="https://code.claude.com/docs/en/how-claude-code-works.md">official docs</a>.</p>
<figure><img src="https://www.imaknas.com/writing/what-survives-compaction/images/table-04.png" width="928" height="346" class="keep-legible" alt="How Codex CLI and Claude Code compact, next to The Crucible before this study."><figcaption>How Codex CLI and Claude Code compact, next to The Crucible before this study.</figcaption></figure>
<p>Codex even warns after compacting that &quot;long threads and multiple compactions can cause the model to be less accurate.&quot; That is Finding 2, stated by its authors.</p>
<p>I added Codex&#x27;s rules to the experiment. Keeping user messages verbatim rescued exactly the user-stated facts: with a cheap summarizer, they went from 17/36 to 36/36, while assistant-stated facts did not move. The cost was 2.3 times more context on every later call.</p>
<h3>The principle: never let the summary be the only copy</h3>
<p>It is tempting to say coding agents and conversational apps simply need different compaction. I think the real distinction is deeper. Every piece of context has three properties:</p>
<ul><li><strong>Authority:</strong> who said it. A user&#x27;s statement is a requirement; the assistant&#x27;s is a claim; a tool&#x27;s is an observation.</li><li><strong>Recoverability:</strong> can it be fetched again? A file can be re-read; a detail said only in chat cannot.</li><li><strong>Liveness:</strong> will it be needed? Current values, decisions and open items will be. Superseded values and dead ends won&#x27;t.</li></ul>
<p>Read that way, Codex&#x27;s design makes sense. User messages are authoritative and unrecoverable, so they stay verbatim. Tool output is recoverable from the repository, so it can go. The summary only has to carry live state, which is why a short handoff is enough. Claude Code&#x27;s CLAUDE.md does the same thing by moving authoritative rules out of the conversation, where compaction can&#x27;t touch them.</p>
<p>Every loss I measured was information whose only copy was the summary. The Crucible never deletes history; it keeps every message in its checkpoint store. The model just couldn&#x27;t see it after pruning. So the fix is not a better summary but making the original reachable again.</p>
<h3>Finding 4: bringing the original back rescues a lossy summary</h3>
<p>I added a recall step. Before each call, it searches the hidden history for the four messages that best match the latest request and quotes them back into context. The search is plain keyword matching weighted by word rarity.</p>
<figure><img src="https://www.imaknas.com/writing/what-survives-compaction/images/table-05.png" width="779" height="168" class="keep-legible" alt="Recall rescued the lossy summaries: 31 → 67 and 9 → 66 of 72 facts."><figcaption>Recall rescued the lossy summaries: 31 → 67 and 9 → 66 of 72 facts.</figcaption></figure>
<p>Recall added 50 and 79 points to the lossy configurations, and changed nothing where the summary already kept everything. The cheapest setup that kept nearly every detail was the cheapest summarizer writing 1k-token handoffs plus recall: 66/72, with each later call reading 2.8k tokens.</p>
<p>Once the original is reachable, the summary no longer has to be a complete record. It only has to be good enough to continue from.</p>
<h3>Finding 5: recall moves the relevance decision to the retriever</h3>
<p>Recall only worked so well because my questions named their subject. So I rewrote every question to describe the subject instead: &quot;Which network port should the firewall allow for the task that clears cached files at the edge?&quot; rather than &quot;Which port does the CDN purge job listen on?&quot; No subject word appears in any of them.</p>
<figure><img src="https://www.imaknas.com/writing/what-survives-compaction/images/table-06.png" width="710" height="273" class="keep-legible" alt="With indirect questions, keyword recall fell from 66 to 28 of 72; letting a model choose the search terms recovered 54."><figcaption>With indirect questions, keyword recall fell from 66 to 28 of 72; letting a model choose the search terms recovered 54.</figcaption></figure>
<p>When the facts are in view, wording doesn&#x27;t matter: the model maps &quot;the task that clears cached files at the edge&quot; to the CDN purge job almost every time. Recall hands that judgment to the retriever, and a keyword or small-embedding retriever is far worse at it than the model. Letting a cheap model read the summary and choose the search terms recovered most of the loss, but not all.</p>
<p>So both designs fail on capability, just in different places. A summary fails in the summarizer, compounding with every pass. Recall fails in the retriever. Whatever decides relevance has to understand the question about as well as the model does.</p>
<h3>A prediction that failed</h3>
<p>I expected facts that sound unimportant when they are said to be dropped more often, even by strong summarizers. So I restated every fact as a throwaway remark (&quot;Side note, probably irrelevant: the staging database listens on port 6543&quot;), changing nothing else.</p>
<p>The opposite happened. A cheap summarizer kept more of the asides: 18/72 instead of 1/72 with Codex-style handoffs. The strong summarizer and recall kept everything either way. In filler where every sentence has the same shape, a hedge makes a sentence stand out rather than fade. Sounding unimportant is not the same as being unimportant to a model, and the harder question, relevance nobody could have foreseen when the fact was stated, remains open.</p>
<h3>Finding 6: summaries drop what the task doesn&#x27;t need yet</h3>
<p>So I tied relevance to the task instead of the wording. Each conversation pursues one stated goal (&quot;this week&#x27;s only priority is the log shipper&quot;). Eight facts are about that goal. Eight more are about another subsystem nobody is working on yet, stated plainly in passing, in exactly the same form as everything else, and mentioned nowhere else. The last turn switches the goal: &quot;Change of plans: the warehouse loader is now the priority.&quot; Then every fact is probed.</p>
<p>Until this point every summarizer I had tested was a Gemini model, so I ran two from each provider family on four conversations:</p>
<figure><img src="https://www.imaknas.com/writing/what-survives-compaction/images/table-07.png" width="669" height="273" class="keep-legible" alt="After a change of plans, OpenAI and Anthropic summaries kept only 4–13 of the 32 off-goal facts; only gemini-3.8-flash kept most (26/32)."><figcaption>After a change of plans, OpenAI and Anthropic summaries kept only 4–13 of the 32 off-goal facts; only gemini-3.8-flash kept most (26/32).</figcaption></figure>
<p>Four of the four OpenAI and Anthropic summarizers lose off-goal facts, and the OpenAI ones keep almost none. Gemini, the family I had been testing all along, does it least.</p>
<p>My first reading, from two Gemini conversations, was that stronger summarizers drop more. It doesn&#x27;t hold across families: the stronger model is more selective within OpenAI and less selective within Anthropic. The family&#x27;s summarizing style matters more than strength. What does hold is the mechanism: a task-focused summary leaves out what the task doesn&#x27;t need yet, and no summarizer can know the plan is about to change.</p>
<p><strong>Recall brings them all back.</strong> Recall looks at the history after the goal has switched, so it should not care about the change of plans. I tested it with the two most goal-selective summarizers:</p>
<figure><img src="https://www.imaknas.com/writing/what-survives-compaction/images/table-08.png" width="799" height="133" class="keep-legible" alt="Recall brought back all 32 off-goal facts for both goal-selective summarizers."><figcaption>Recall brought back all 32 off-goal facts for both goal-selective summarizers.</figcaption></figure>
<p>Deciding at read time recovers exactly what deciding at write time loses. This is the one result in the study I expect to survive new models: a summary is written for the goal of the moment, no summarizer can anticipate a change of plans, and only a design that keeps the originals reachable is immune. One caveat carries over from Finding 5: these questions named their subject, so keyword recall was enough. With the same questions asked indirectly, keyword recall restored only 16–17 of the 32 and embedding recall did no better, while model-guided recall, where a small model reads what is still visible (including the turn that changed the goal) and chooses what to fetch, restored 28–30. The two summarizers here are two replications, not a comparison: they differ in family and tier.</p>
<p><strong>Choosing what to fetch doesn&#x27;t need a strong model.</strong> Until then the model choosing the search terms had always been gpt-6-luna, the cheapest in the app. So I fixed one summarizer (gpt-6-sol, the most goal-selective), wrote its summaries once, and replayed them with six different models choosing what to fetch, a cheap and a strong one from each family. Same indirect questions, four conversations:</p>
<figure><img src="https://www.imaknas.com/writing/what-survives-compaction/images/table-09.png" width="622" height="308" class="keep-legible" alt="Every model choosing what to fetch beat keyword recall; the cheapest, gpt-6-luna (30/32), matched the strongest."><figcaption>Every model choosing what to fetch beat keyword recall; the cheapest, gpt-6-luna (30/32), matched the strongest.</figcaption></figure>
<p>Every model beats keyword recall. Luna and sol are statistically equivalent, and the strong models are equivalent to luna; only the smallest Gemini falls somewhat behind. The cheapest model made the call as well as gemini-3.8-flash, which cost 24 times as much for the same calls ($0.59 against $0.024). That fits the time-of-decision story: before the question is known, no summarizer can judge relevance reliably; once it is known, even a small model can. Which cheap model is good enough will change with every release. That deciding late makes the decision easy should not.</p>
<p><strong>Why the summary drops them.</strong> &quot;No summarizer can anticipate a change of plans&quot; is a claim, so I tested it. I reran the same conversations with one change: every goal reminder also said &quot;once that is done, the next priority will be X&quot;. The facts, filler and positions stayed identical. Separately, I tried a generic hedge that knows nothing about what comes next: the usual instruction plus &quot;priorities may change; keep the specific values for every subject, not only the current goal.&quot;</p>
<figure><img src="https://www.imaknas.com/writing/what-survives-compaction/images/table-10.png" width="711" height="203" class="keep-legible" alt="Announcing the next goal took gpt-6-sol from 4 to 30 of 32; the generic hedge worked only with 9× longer summaries."><figcaption>Announcing the next goal took gpt-6-sol from 4 to 30 of 32; the generic hedge worked only with 9× longer summaries.</figcaption></figure>
<p>For gpt-6-sol, knowing the next goal is the whole story: 4 to 30 of 32, for summaries 30% longer. The generic hedge also kept everything, but only by barely compressing: summaries of 5.5k tokens against an 8k threshold, rewritten twice as often. Without knowing the future, the only way to keep off-goal facts at write time is to give up the compression. Claude Haiku shows a second, separate loss: told the future and hedged, it still kept only 20 of 32, and it lost on-goal facts too. That looks like fidelity, not selection. Read-time recall sidesteps both, because the originals are still there. Was it only that the subject&#x27;s name kept coming up? Announcing the next goal names it a dozen times, so I ran a control that names it just as often, as out of scope this quarter, with the change of plans still a surprise. gpt-6-sol kept 6 of 32: about 2 of the 26 recovered facts come from the name standing out, and the other 24 from knowing it comes next.</p>
<h3>Finding 7: once the originals are reachable, the cheapest summarizer is enough</h3>
<p>Recall changes what a summarizer is for, so I chose the app&#x27;s default summarizer again, this time with recall on, as the app runs. Before running, I wrote down the rule: keep the candidates statistically no worse than the best on both scenarios, then take the cheapest per summary at the prices in force from January 2027 (gemini-3.8-flash&#x27;s introductory price ends in December). Four summarizers, the detail-heavy scenario and the change-of-plans scenario, four conversations each, strict oracle:</p>
<figure><img src="https://www.imaknas.com/writing/what-survives-compaction/images/table-11.png" width="928" height="221" class="keep-legible" alt="With recall on, gpt-6-luna's summaries matched the best at about 1/25 of the cost; alone they keep much less."><figcaption>With recall on, gpt-6-luna's summaries matched the best at about 1/25 of the cost; alone they keep much less.</figcaption></figure>
<p>Only gpt-6-luna passed the rule: statistically equivalent to the best on the detail-heavy scenario, itself the best after a change of plans, and about 1/25 of gemini-3.8-flash&#x27;s cost per summary (costs are for the detail-heavy scenario at 2027 prices). Alone, its summaries keep less than half as much. Once the originals can be fetched, the summary&#x27;s job shrinks from being the record to being a pointer, and a pointer can be cheap. The choice holds only because recall is on: without it, the ranking inverts.</p>
<p>One result I can&#x27;t explain yet: recall added 16 to 29 facts for every other summarizer, but nothing for gemini-3.8-flash after the change of plans (56 either way), and a later audit reproduced the gap.</p>
<h3>Keeping the scaffolding honest</h3>
<p>Every mechanism in this post, from the summary itself to the verbatim window and recall, exists because it beat a baseline on the models I tested. A newer model can close that gap on its own, and scaffolding that no longer helps can get in its way. So each mechanism on the app&#x27;s default path now declares its baseline and its claim, and an audit re-measures the pair on current models whenever the model list changes, recording keep, retire, harmful or inconclusive.</p>
<p>The first two audits already changed the picture. Both used two cheap answering models on the detail-heavy scenario:</p>
<figure><img src="https://www.imaknas.com/writing/what-survives-compaction/images/table-12.png" width="573" height="133" class="keep-legible" alt="Recall went from a mixed blessing (first audit, thorough summaries) to essential (second audit, cheap summaries)."><figcaption>Recall went from a mixed blessing (first audit, thorough summaries) to essential (second audit, cheap summaries).</figcaption></figure>
<p>With a thorough summary, the recalled excerpts were a distraction for the weaker model: it took values from the wrong subject and gave the same answer to different questions. With a thin summary, the excerpts are the main source, and recall became essential for both. The second audit also checked the summarizer choice with real answering models: neither pricier summarizer did better for either reader.</p>
<p>Not every verdict should be acted on. The audit flagged the verbatim window, which keeps the newest messages as they are, as &quot;retire&quot;: with recall on, it added nothing to fact recall. But the probes only ask about planted facts, and the window exists for continuity with the last few turns, which they don&#x27;t measure. An audit is only as good as its claim.</p>
<h3>What changed in the app</h3>
<ul><li>The re-summarize loop is fixed, and prompt caching works on all three providers.</li><li>Recall is on by default: once a thread has been summarized, the cheapest available model picks which original messages the new request is about and shows them again. It costs nothing until a thread compacts, then one small call per turn.</li><li>The default summarizer is gpt-6-luna, chosen above, and any other can be picked in the Control Panel.</li><li>Every mechanism on the default path is re-audited when the models change.</li><li>State-and-index summaries, keyword recall and keeping user messages verbatim are measured here but are not defaults.</li></ul>
<h3>What it cost, and what went wrong</h3>
<p>The whole study cost roughly $80 in API calls. About $41 of that went on the first four runs, before I had a budget guard. The rest, about twenty more experiments, cost about $36 under hard caps, plus about $3 that a cost estimate spent by mistake, including one run that returned nothing and a few dollars of calibration for runs that were then refused.</p>
<p>The mistakes were as instructive as the results:</p>
<ul><li><strong>A ceiling effect.</strong> The first pilot scored 100% everywhere because the planted facts were the only specific content in the conversation.</li><li><strong>A scoring bug.</strong> Models often answer and then quote their source, such as &quot;Litware. It said: &#x27;we picked Litware over Margie and Contoso&#x27;&quot;. Matching anywhere in the reply counted the quoted rejected vendors as wrong. Scoring now reads only the committed first line.</li><li><strong>The re-summarize loop</strong> from Finding 3, found only because the first expensive run wrote 132 summaries.</li><li><strong>Guessed prices.</strong> I assumed a summarizer cost $2.50 per million output tokens; it was $9. I had also picked &quot;nano&quot; models as the cheap default by name, while a newer model cost half as much. Prices now live in a table with a source for each, and experiment defaults are chosen from it.</li><li><strong>A strict oracle that was too strict.</strong> My first &quot;value next to its subject&quot; check used a 200-character window. Summaries organized in long per-topic sections broke it, while every model still answered correctly.</li><li><strong>A noisy estimator.</strong> One calibration call per model is not enough. One run cost 1.8 times its estimate. The hard cap, not the estimate, is what bounds the spend.</li></ul>
<p><strong>A capped run that returned nothing.</strong> A long-horizon run (about 22 summaries per conversation) spent its whole $5 cap and produced zero results. Every condition was replaying at once, the summaries grew with every fold faster than my one-call calibration predicted, and the cap stopped all of them mid-way. Runs now go one condition at a time, cheapest first, each admitted only if the remaining budget covers it and saved as soon as it finishes.</p>
<p><strong>An estimate that spent money.</strong> The cost estimate is a dry run with stand-in models, so it should be free. But the model choosing what to fetch was handed the real model client when the experiment was set up, and the dry run never swapped it out. Estimating the six-model run above made about 384 real calls, roughly $3, recorded nowhere and priced at zero. Now each step that calls a model says which model it calls, and the estimate meters it with a stand-in.</p>
<p><strong>The cache that never hit.</strong> Cache hits across the experiments were about 1%. The app&#x27;s system prompt began with the current time to the second, so no two requests shared a prefix. Moving the time out of the system prompt was not enough, because each provider caches differently:</p>
<figure><img src="https://www.imaknas.com/writing/what-survives-compaction/images/table-13.png" width="928" height="276" class="keep-legible" alt="Where per-request context has to go for prompt caching to work, by provider."><figcaption>Where per-request context has to go for prompt caching to work, by provider.</figcaption></figure>
<p>The general rule: anything that changes per request (a timestamp, retrieved excerpts) belongs after the last point the next request will repeat, and where that point is depends on the provider.</p>
<h3>Related work</h3>
<p>These results sit next to a growing body of work, and several papers point the same way:</p>
<ul><li><a href="https://arxiv.org/abs/2508.21433">The Complexity Trap</a> (JetBrains Research, NeurIPS 2025 DL4Code workshop): in coding agents, simply masking old tool output matches LLM summarization at about half the cost.</li><li><a href="https://factory.ai/news/evaluating-compression">Evaluating Context Compression</a> (Factory.ai): probe-based evaluation of compaction in real agent sessions. It credits merging each summary into a persistent state instead of regenerating it, which lines up with the state-and-index result here.</li><li><a href="https://arxiv.org/abs/2609.26779">CliffCompaction</a> (Nguyen, Cho, Chen, Dettmers, 2026): never compacts already-compacted content, to avoid drift. That is the per-pass loss measured in Finding 2.</li><li><a href="https://arxiv.org/abs/2608.06503">Toward Reliable Context Compression for Long-Horizon Agents</a> (2026): repeated compression makes agent runs unstable.</li><li><a href="https://arxiv.org/abs/2608.16370">What Does Context Compression Cost an Agent?</a> (2026): completion rates stay flat while agents re-fetch dropped state, so the cost hides in extra calls.</li><li><a href="https://arxiv.org/abs/2601.00821">Fidelity Before Structure</a> (2026): for conversation memory, retrieving verbatim chunks beats retrieving extracted artifacts.</li><li>Background: <a href="https://arxiv.org/abs/2308.15022">recursive summarization for dialogue memory</a>, Letta/MemGPT&#x27;s <a href="https://vectorize.io/articles/mem0-vs-letta">tiered memory with searchable recall</a>, Anthropic&#x27;s <a href="https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents">context engineering guide</a>, and Chroma&#x27;s <a href="https://www.trychroma.com/research/context-rot">Context Rot</a>.</li></ul>
<p>What I haven&#x27;t seen elsewhere, though I may have missed it: a controlled once-versus-incremental comparison to separate per-pass loss from capacity, an oracle that separates &quot;the summary dropped it&quot; from &quot;the model missed it&quot;, a paired test of repetition, a causal test of why task-focused summaries drop facts (announcing the next goal versus only naming it), and an audit that re-checks each piece of context machinery against going without it whenever the models change.</p>
<h3>Limitations and what&#x27;s next</h3>
<p>This is a pilot study, not a benchmark:</p>
<ul><li><strong>Synthetic conversations.</strong> Real conversations are messier, and the next step is to replay real threads.</li><li><strong>Small samples.</strong> Each cell is 32 to 96 paired probes from three or four conversations, with small and mid-size models only.</li><li><strong>One threshold and no multiple-comparison correction.</strong> Treat single comparisons as indicative.</li><li><strong>Task relevance is preliminary.</strong> Findings 6 and 7 rest on four conversations per condition, scored by the oracle. Real answering models were checked only on the detail-heavy scenario, in the audits, not after a change of plans. Real conversations are still to come.</li><li><strong>String-match scoring.</strong> It is strict, and it only reads a reply&#x27;s first line.</li></ul>
<p>Next I want to run it on real conversation histories, and add stronger models. The goal is to turn this into a small open-source context-management layer that decides what to keep verbatim, what to summarize and what to leave retrievable.</p>
<p><strong>Code and data.</strong> Everything runs in <a href="https://github.com/imaknas/the-crucible">The Crucible</a> (<code>crucible compaction-eval</code> and <code>crucible scaffold-audit</code>). Every number here is in the <a href="https://github.com/imaknas/the-crucible/blob/main/backend/experiments/NOTES.md">lab notes</a>, with the raw results next to them.</p>
<p><strong>How this was made.</strong> I ran this study with Claude Code as a coding and analysis assistant. It wrote much of the experiment code and the first draft of this post. I directed the questions and checked the results, and I&#x27;m responsible for the conclusions.</p>]]></content:encoded>
  </item>
  <item>
    <title>Nobody made a decision. We just described</title>
    <link>https://www.imaknas.com/writing/nobody-made-a-decision/?lang=en</link>
    <guid isPermaLink="false">https://www.imaknas.com/writing/nobody-made-a-decision/#en</guid>
    <pubDate>Sun, 24 May 2026 12:00:00 +0000</pubDate>
    <description>A no-AI spec challenge with my team showed that most engineers weren&#x27;t making decisions, only describing the system.</description>
    <content:encoded><![CDATA[<p><em>What a no-AI spec challenge revealed about how engineers actually think.</em></p><figure><img alt="" src="https://www.imaknas.com/writing/nobody-made-a-decision/images/en-1.png" width="960" height="540"></figure><p>Last week I ran a competition with my team.</p><p>The rules were simple: one ambiguous requirement, 30 minutes without AI, write a spec. Then hand the spec to <strong>Claude Code</strong> and let it build. At the end, everyone presents. The team votes on the best output.</p><p>I expected the results to tell me who was good at specs. They told me something more revealing than that.</p><h2>The Setup</h2><figure><img alt="" src="https://www.imaknas.com/writing/nobody-made-a-decision/images/en-2.jpeg" data-zoom="/writing/nobody-made-a-decision/images/en-2@2x.jpg" width="540" height="720"></figure><p>The requirement was deliberately vague — modeled closely on work we actually do:</p><blockquote>When a data owner publishes a dataset listing and a buyer completes a purchase, the system must deduct payment from the buyer’s wallet, generate a transaction record, update both wallet states, and notify relevant parties. Consider idempotency and failure compensation.</blockquote><p>Eight engineers. Individual work. No AI during the spec phase. After that, <strong>Claude Code</strong> only.</p><p>I’d spent the previous two sessions teaching the team about <strong>Harness Engineering</strong> — the idea that the environment you give an AI matters as much as the prompt. That specs are how you engineer that environment. That the bottleneck isn’t the model.</p><p>This was the test.</p><h2>What I Saw</h2><p>Some people had nothing to start from.</p><p>A few specs were nearly verbatim copies of the requirement. The steps section said: <em>deduct payment, generate transaction record, update wallet states, notify relevant parties</em>. Exactly what the requirement said. No decomposition. No sequencing. No edge cases.</p><p>This wasn’t laziness. It was something more fundamental: there was no mental model underneath to translate the requirement into structure. When you take away the tool that would normally help build that structure, there’s nothing left to write.</p><h2>Deep domain knowledge didn’t help as much as I expected.</h2><p>One of the most experienced engineers on the team — someone who had spent years working on payment systems — got stuck. Their verbal feedback during the exercise was: the requirement wasn’t defined clearly enough. They needed more context before they could design.</p><p>That’s how experienced engineers used to work. Someone writes a PRD. You design against it. If the PRD is unclear, you push back and ask for clarification.</p><p>But the challenge wasn’t asking them to design against a complete requirement. It was asking them to <em>resolve the ambiguity themselves</em> — to treat the fuzzy requirement as the problem to be structured, not a blocker to complain about. That’s a different skill. Years of domain knowledge don’t automatically transfer to it.</p><h2>The winner handed full control to AI and stepped away.</h2><p>One engineer opened <strong>Claude Code</strong> in auto mode, let it run, and came back to results. Their spec wasn’t the most precise. But it was structured well enough that the AI produced something that looked complete and convincing. The team voted it the best output.</p><p>The delegation itself wasn’t the problem. The problem surfaced at the end, when the output came back.</p><p>Most people said the AI added things they hadn’t specified — saga patterns, retry logic, compensating events. And most people admitted they weren’t sure, in the time they had, whether those additions were right or wrong.</p><p>That’s the real gap. Not that the spec was incomplete — every spec is incomplete. It’s that when the AI makes a decision on your behalf, you need a prior decision of your own to evaluate it against. If you never decided how failure compensation should work, you have no ground to stand on when the AI proposes a saga. You can accept it or reject it, but either way you’re guessing.</p><p>A spec that captures your decisions doesn’t prevent the AI from adding things. It gives you the ability to judge what it adds.</p><h2>The Pattern Nobody Told Me About</h2><figure><img alt="" src="https://www.imaknas.com/writing/nobody-made-a-decision/images/en-3.jpeg" data-zoom="/writing/nobody-made-a-decision/images/en-3@2x.jpg" width="960" height="720"></figure><p>After the spec phase, I asked everyone to share one thing: <em>what was the most important decision you made in your spec?</em></p><p>Nobody answered the question. At first I thought they were avoiding it. Then I realized I had never given them a reason to think in decisions in the first place. The spec template I provided had structure — trigger, steps, edge cases — but no choice points. I asked them to describe a system, then expected them to have made decisions about it. That’s not fair.</p><h2>Description versus decision</h2><p>A decision sounds like: <em>“I chose to handle failure at the wallet debit step separately from the credit step, because if debit succeeds and credit fails, the rollback logic is different — and I didn’t want the AI to collapse those into a single failure handler.”</em></p><p>One person got close. Their verbal share included: if debit succeeds but credit fails, retry three times; if retry fails, emit a compensating event and close the transaction. That’s a real decision — they had considered multiple options and chosen one with a reason.</p><p>Everyone else described their spec. The description was sometimes accurate. But it wasn’t a decision.</p><p>When I reflected on this afterward, I realized the problem wasn’t that they couldn’t articulate decisions. It was that <em>they hadn’t made any.</em> Spec writing, for most of them, meant transcribing their first intuition about how the system should work — not choosing between alternatives.</p><h2>Why This Matters More Than It Used To</h2><p>In the old workflow, the absence of explicit decisions was survivable. You wrote code. Someone reviewed it. The decision got surfaced through the review cycle. You had time to course-correct.</p><p>AI-assisted development removes that forcing function.</p><p>When <strong>Claude Code</strong> executes your spec in minutes, every assumption you didn’t make explicit becomes a decision the AI made for you. It will fill the gaps confidently. It will produce something that looks complete. And you won’t know which gaps it filled until you’re debugging the result.</p><p>The spec isn’t just a planning document anymore. It’s the artifact where your decisions live. If you didn’t make decisions in it, you’re not directing the AI — you’re letting the AI direct itself, and signing off on the output.</p><h2>What I’m Changing for the Next Session</h2><p>The exercise revealed a gap I didn’t know how to see before: the difference between <em>describing a system</em> and <em>making decisions about a system</em>.</p><p>Next time, the spec template will include explicit choice points. Not just “what are your steps” — but “if debit succeeds and credit fails, which do you choose: retry, rollback, or compensating event? Why?” Forcing them to choose between named alternatives is the only way I know to get people into decision mode rather than description mode.</p><p>The voting format is changing too. Peer voting on overall output tends to favor output that looks complete over output that is precise. Next time I’m adding a scoring rubric tied to the three things that actually matter: spec coverage, alignment between spec and output, and skeleton clarity.</p><p>And before we start, I’ll spend ten minutes showing them what I found. Not to call anyone out — but because the most clarifying insight from this exercise is also the most useful one:</p><p><em>Most of the team didn’t know they weren’t making decisions. They thought describing the system was the same thing.</em></p><p>That’s the gap AI exposes. Not skill. Not effort. The assumption that your first intuition, written down, is a spec.</p><figure><img alt="" src="https://www.imaknas.com/writing/nobody-made-a-decision/images/en-4.jpeg" data-zoom="/writing/nobody-made-a-decision/images/en-4@2x.jpg" width="960" height="720"></figure><p><em>I run AI engineering workshops for software teams navigating this transition. If your team is working through the same shift, I’d like to hear about it — cshiauknas@gmail.com</em></p>]]></content:encoded>
  </item>
  <item>
    <title>AI Raises the Cost of Not Thinking</title>
    <link>https://www.imaknas.com/writing/ai-raises-the-cost-of-not-thinking/?lang=en</link>
    <guid isPermaLink="false">https://www.imaknas.com/writing/ai-raises-the-cost-of-not-thinking/#en</guid>
    <pubDate>Wed, 13 May 2026 12:00:00 +0000</pubDate>
    <description>Extracting adaptive-iteration from a self-running content pipeline, and why AI makes unclear thinking more expensive, not less.</description>
    <content:encoded><![CDATA[<figure><img alt="" src="https://www.imaknas.com/writing/ai-raises-the-cost-of-not-thinking/images/en-1.jpg" data-zoom="/writing/ai-raises-the-cost-of-not-thinking/images/en-1@2x.jpg" width="1321" height="720"></figure><p>In March 2026, Andrej Karpathy published a project called <a href="https://github.com/karpathy/autoresearch">autoresearch</a>. The idea: give an AI agent a real LLM training setup and let it experiment overnight. It modifies code, trains for five minutes, checks if the result improved, keeps or discards, and repeats. You wake up to a log of experiments and a better model.</p><p>His framing of the human’s role stuck with me: <em>“You’re not touching any of the Python files like you normally would as a researcher. Instead, you are programming the program.”</em></p><p>That’s the shift. Not writing code. Programming the system that writes the code.</p><p>I’d been living a version of this for weeks — not in ML research, but in content creation. Every day, an AI agent running on <a href="https://openclaw.ai/">OpenClaw</a> generates scripts, produces videos, publishes them, pulls analytics, and runs A/B experiments on what worked. I’m not in the loop for any of that execution. What I do is define what the system should be optimizing for, catch where it’s going wrong, and redirect it before it drifts.</p><p>At some point I realized the pattern underneath all of this had nothing to do with video. It was the same loop Karpathy was describing — just in a different domain. So I extracted it into an open-source framework: adaptive-iteration.</p><p>But the more interesting story isn’t what we built. It’s what building it revealed.</p><h2>The framework wasn’t designed. It emerged.</h2><p>adaptive-iteration describes a four-part loop: produce something, measure it, generate hypotheses about what to change, then periodically challenge whether you&#39;re optimizing for the right thing at all. Repeat.</p><p>Here&#39;s what&#39;s strange: the framework itself was built through exactly that process.</p><p>It wasn&#39;t designed upfront from first principles. It emerged from running a real system — a content pipeline that had been iterating on itself for weeks — and recognizing a structure that kept reappearing under the domain-specific noise. I described that structure to an AI agent running on OpenClaw. We built the abstraction together. I defined what it should be. The agent implemented it. I caught what was wrong — a dependency direction that was backwards, implementation details that didn&#39;t belong in a public repo, architectural decisions that needed to be made explicit. The agent fixed it. That cycle repeated until something clean enough to ship emerged.</p><p>The tool was made by the process it describes. That&#39;s not a coincidence — it&#39;s a signal. The best abstractions don&#39;t come from design. They come from running something real long enough to see what&#39;s actually there.</p><p>Karpathy&#39;s autoresearch is a specific instantiation of this loop, hardwired to ML training. adaptive-iteration is the same loop made portable — a Ledger that records what happened, an Analyzer that surfaces what worked, a HypothesisEngine that reasons about what to try next, and a DomainAdapter interface that&#39;s the only layer touching your actual system. Bring your own domain. The framework handles the rest.</p><h2>AI doesn’t lower the cost of building. It raises the cost of not thinking clearly.</h2><p>Here’s the part the “AI democratizes everything” narrative gets wrong.</p><p>Early in the process, I told the agent to push the framework to GitHub. It did. But when I looked at what was actually pushed, one of the adapter files wasn’t a clean example — it was the real implementation, with hardcoded private paths, directly exposing the internals of a production system I hadn’t intended to share. The agent had done exactly what was asked. The problem was what was asked hadn’t been thought through clearly enough.</p><p>In the old world, this kind of mistake had friction built in. Building was slow. A fuzzy mental model had time to get clarified during implementation — the cost of the process forced you to think. When the feedback loop is instant, that forcing function disappears. The agent executes your mental model at full speed, including its flaws.</p><p>This is the thing nobody says: AI doesn’t make building easier. It makes the quality of your mental model more consequential, faster. A clear model ships something good overnight. A fuzzy model ships something wrong overnight — and it’s already on PyPI before you notice.</p><p>The bottleneck didn’t disappear. It transformed. And it’s less forgiving than the old one.</p><h2>The governance asymmetry</h2><p>Which brings us to the part of the paradigm shift that’s less comfortable to say out loud.</p><p>“Anyone can build now” is technically true. But what most of these conversations miss is that building and building <em>the right thing</em> are different problems. The first is an execution problem. The second is a thinking problem. AI solves the first. It amplifies the second.</p><p>The new bottleneck — recognizing generalizable patterns across domains, holding a clear mental model of what a system should be, knowing which architectural decisions matter and which don’t — isn’t more democratically distributed than the ability to code. It’s just different. A person who can recognize that their content optimization loop and their proposal strategy loop share the same underlying structure, and can govern an AI agent to extract that abstraction cleanly, will compound enormously from these tools. Someone who can’t will build faster and make more mistakes faster.</p><p>What’s genuinely democratized is the ceiling. A researcher who has spent years in a domain but never learned to code can now build the analysis tools their field has needed. A teacher can build adaptive curriculum systems. A founder can build the product they’ve been describing to developers who never quite got it right. The ideas that stayed ideas because the person who had them couldn’t implement them — those ideas are now executable.</p><p>But the floor didn’t move. Governance — the ability to define what should be built, evaluate what was built, and catch the gap between the two — is still a skill. It still has to be developed. Karpathy’s autoresearch works because Karpathy is Karpathy. The program.md he writes to direct his agents carries decades of ML research intuition. The output quality is a function of the governance quality.</p><p>That’s the paradigm shift I actually care about. Not that execution is cheap. That governance is now the core competency — and unlike code, nobody’s teaching it yet.</p><p>Three questions I deliberately left open in this piece.</p><p>What happens in domains where “Measure” is the hard problem — where feedback cycles are long, noisy, or fundamentally contested? The framework assumes you can observe outcomes. A lot of the highest-value decisions in the world don’t come with a loss function.</p><p>What does governance actually decompose into as a skill? I called it a competency and then left it there. That’s not good enough. The difference between intuition and a teachable framework is whether you can name the components.</p><p>And the one that keeps me up at night: AI is already pushing into the governance layer — LLM-as-a-Judge, agentic workflows with self-correction, systems that generate their own hypotheses. The boundary I described between human governance and AI execution is already moving. What do you govern when the thing you’re governing is itself governing something else?</p><p>I built a framework that claims to be domain-agnostic. I’m not sure I’ve tested that claim honestly. I’m still working on all three.</p><h2>The repo</h2><p>adaptive-iteration is a domain-agnostic experimentation framework. Bring your own domain — the framework handles the experiment → measure → learn → challenge loop.</p><p>→ <a href="https://github.com/imaknas/adaptive-iteration">github.com/imaknas/adaptive-iteration</a><br>→ <a href="https://pypi.org/project/adaptive-iteration">pypi.org/project/adaptive-iteration</a></p>]]></content:encoded>
  </item>
  <item>
    <title>Why Your Team Is Using AI Wrong — And How One Session Changed That</title>
    <link>https://www.imaknas.com/writing/why-your-team-is-using-ai-wrong/?lang=en</link>
    <guid isPermaLink="false">https://www.imaknas.com/writing/why-your-team-is-using-ai-wrong/#en</guid>
    <pubDate>Sun, 10 May 2026 12:00:00 +0000</pubDate>
    <description>An internal AI coding session: the missing skill isn&#x27;t prompting, it&#x27;s harness engineering.</description>
    <content:encoded><![CDATA[<figure><img alt="" src="https://www.imaknas.com/writing/why-your-team-is-using-ai-wrong/images/en-1.png" width="1371" height="720"></figure><p><em>The missing skill isn’t prompting. It’s Harness Engineering.</em></p><p>Last week I ran an AI coding session at my company, walking the team through how I actually use Claude Code day to day.</p><p>What I discovered surprised me: the gap between people getting 10x productivity and people feeling like AI is overhyped isn’t about skill, intelligence, or even experience with AI tools. It’s about <strong>exposure to a good working model</strong>.</p><p>Most people have never seen what high-efficiency AI-assisted coding looks like. So they don’t know what they’re missing.</p><h2>The concept: Harness Engineering</h2><p>The industry is converging on a term for this: harness engineering. Both <a href="https://www.anthropic.com/engineering/effective-harnesses-for-long-running-agents">Anthropic</a> and <a href="https://openai.com/index/harness-engineering/">OpenAI</a> engineering teams have published on the practice — the core idea being that engineering the <em>environment</em> an agent operates in matters as much as the prompts you give it. The insight is simple but profound: a capable model running in a poorly designed environment will consistently underperform. The harness — the initialization, the constraints, the structured context — determines whether your agent stays on task or goes off the rails. Most developers skip this entirely. They open a chat window and start prompting.</p><h2>The workflow I showed them has three steps.</h2><h2>Step 1: /init — build the harness first</h2><p>Before writing a single line of code, run /init in Claude Code. This is your harness initialization: the agent reads your project structure, tech constraints, and coding conventions, and creates a shared context that every subsequent session inherits. This is the difference between an agent that understands <em>your</em> codebase and one that&#39;s making educated guesses. Every engineer on the team working with the same initialized harness gets consistent, codebase-aware output. Skip this, and you&#39;re not doing AI-assisted development — you&#39;re doing expensive autocomplete.</p><h2>Step 2: ticket → spec — clarify before you execute</h2><p>A Jira ticket tells you <em>what</em> to build. A spec tells the agent <em>how to know when it’s done</em>. Before any implementation, work with the agent to convert the ticket into a proper spec: acceptance criteria, edge cases, integration points. During this process, the agent surfaces gaps in the requirements — things you hadn’t fully thought through. This isn’t extra work. It’s the planning that should have happened anyway, now with an agent that actively helps you find the holes. Anthropic’s research found that decomposing work into clearly-scoped units — rather than one-shotting a full feature — was one of the biggest levers for reliable agent performance. The spec is how you create those units.</p><h2>Step 3: review the plan, then let go</h2><p>Once the spec is solid, the agent generates an implementation plan. Review the plan — not the code, just the direction. Is the approach right? Does it account for what you care about? When you’re satisfied, let it run. Human sets direction. Agent handles execution. This maps directly to what Anthropic calls the “coding agent” pattern: give the agent a clear scope, let it make incremental progress, check back when it’s done.</p><h2>The real insight</h2><p>AI coding tools don’t underperform because the AI isn’t good enough. They underperform because we hand them vague inputs and expect precise outputs. The engineers on my team who already had strong spec discipline picked this up immediately. Those used to “figuring it out as they go” found it harder — not because of the AI, but because it exposed gaps in how explicitly they were thinking about requirements. That’s the uncomfortable truth about AI productivity: it amplifies your existing engineering habits. Good habits scale. Vague habits stay vague, just faster.</p><h2>If you’re leading a team using AI coding tools</h2><p>Don’t just give people access and hope for the best. Run a session. Show them what a properly initialized harness looks like. Let them see the difference between dropping a ticket into Claude and building a spec together first. The bottleneck isn’t the model. It’s whether your team knows how to engineer the environment it runs in.</p>]]></content:encoded>
  </item>
  <item>
    <title>From Frontend/Backend to a Spectrum: Drawing a Clear Map for the Chaotic AI Talent Market</title>
    <link>https://www.imaknas.com/writing/ai-job-spectrum/?lang=en</link>
    <guid isPermaLink="false">https://www.imaknas.com/writing/ai-job-spectrum/#en</guid>
    <pubDate>Sat, 09 Aug 2025 12:00:00 +0000</pubDate>
    <description>Borrowing the frontend/backend split to map muddled AI job titles onto a spectrum from model research to system building.</description>
    <content:encoded><![CDATA[<p>The age of AI has arrived, but the job market seems to be lost in a fog. The “Machine Learning Engineer” at Company A does the work of the “AI Engineer” at Company B, while the “Data Scientist” at Company C might not touch models at all. We lack a common language to describe the different professional functions in the age of AI, making it difficult for companies to hire and for talent to position themselves.</p><p>The purpose of this article is not to invent more new titles, but to provide a clear analytical framework to help individuals and companies better understand and define their value in the AI wave. It is important to note that this classification is not an industry standard, but rather a summary based on my personal observation and experience, aimed at helping readers quickly find their own position and that of their team.</p><h3>A Familiar Analogy: Understanding AI Roles Through “Frontend” and “Backend”</h3><p>To simplify this complex issue, we can borrow the most classic model from software development, the frontend/backend separation, to make an analogy.</p><ul><li><strong>The AI Frontend (The “Model Layer”):</strong> The core work revolves around the “model itself.” Its mission is to research, train, and optimize algorithms, pursuing improvements in the model’s intrinsic metrics. The corresponding roles are scientists and algorithm experts.</li><li><strong>The AI Backend (The “System Layer”):</strong> The core work is to take a model as a component to build a stable, scalable, and large-scale application system that solves real business problems. Its mission is to ensure the reliability and ultimate business value of the entire AI application.</li></ul><figure><img alt="AI frontend and AI backend: one revolves around the model itself, the other takes it as a component" src="https://www.imaknas.com/writing/ai-job-spectrum/images/layers-en.png" width="696" height="398" class="keep-legible"><figcaption>AI frontend and AI backend: one revolves around the model itself, the other takes it as a component</figcaption></figure><p>This simple dichotomy helps us make the first crucial distinction, allowing us to see two vastly different value-creation paths in the AI domain.</p><h3>The Model’s Limits: When the Frontend/Backend Line Blurs</h3><p>However, this concise model reveals its limitations when faced with the complex tasks of the real world. For example, a critical, high-value job: <strong>model fine-tuning</strong>. Where does it belong?</p><ul><li>From the perspective of business purpose and data source, it serves a specific application, belonging to the backend.</li><li>From the perspective of required skills and execution process, it demands deep model knowledge, belonging to the frontend.</li></ul><p>Fine-tuning holds a unique position on the spectrum. As it involves both adjusting model parameters and considering practical application scenarios, I see it as a key bridge role between the “model side” and the “system side.”</p><figure><img alt="Fine-tuning spans the model side and the system side" src="https://www.imaknas.com/writing/ai-job-spectrum/images/bridge-en.png" width="696" height="278" class="keep-legible"><figcaption>Fine-tuning spans the model side and the system side</figcaption></figure><h3>A More Precise Map: The “AI Job Spectrum” Model</h3><p>A more accurate perspective is to view AI functions as a continuous spectrum. It clearly delineates the value-creation process at every step, from pure scientific research to the final business application. To compare these roles within the same context, I will place them on a spectrum that runs from “model research” to “system implementation.”</p><figure class="figure-wide"><img alt="AI job spectrum table: from model research to system implementation, comparing Research Scientist, Applied Scientist, Data/ML Scientist, Machine Learning Engineer, AI Systems Architect, AI Engineer and AI Consultant/SA by spectrum position, core task, value measured and tech stack." src="https://www.imaknas.com/writing/ai-job-spectrum/images/en-1.png" width="1200" height="675" class="keep-legible"><figcaption>AI Job Spectrum</figcaption></figure><p>This more granular spectrum clearly delineates the difference between an “Applied Scientist” and a “Data Scientist”: the former is a <strong>pioneer</strong> exploring unknown possibilities, while the latter is a <strong>master craftsman</strong> who uses mature technologies to create stable value for businesses.</p><p>More importantly, it unpacks the most critical engineering zone between “model” and “product.” It is no longer a vague “center,” but an <strong>“engineering trilogy”</strong> composed of three key roles:</p><ul><li>The <strong>Machine Learning Engineer (MLE)</strong> turns models into stable, callable services.</li><li>The <strong>AI Systems Architect</strong> designs a grand application blueprint using these services and other system components.</li><li>The <strong>AI Engineer</strong> implements the final product features according to that blueprint.</li></ul><figure><img alt="The engineering trilogy: from model to product features" src="https://www.imaknas.com/writing/ai-job-spectrum/images/trilogy-en.png" width="696" height="261" class="keep-legible"><figcaption>The engineering trilogy: from model to product features</figcaption></figure><p>A typical output from an AI Systems Architect is not just a technical specification document, but a complete business and technical battle plan. For example, it would clearly define a phased implementation strategy:</p><ul><li><strong>Phase 1:</strong> Focus on building a stable, secure, and observable private AI infrastructure (e.g., a production-grade RAG system).</li><li><strong>Phase 2:</strong> On this foundation, plan high-value applications that generate significant business returns (e.g., personalized data analysis, enterprise-grade intelligent search).</li><li><strong>Phase 3:</strong> Ultimately, through more advanced technologies (e.g., domain-specific model fine-tuning), build a deep technical moat that competitors cannot easily surpass.</li></ul><figure><img alt="An AI Systems Architect’s phased implementation strategy" src="https://www.imaknas.com/writing/ai-job-spectrum/images/phases-en.png" width="696" height="294" class="keep-legible"><figcaption>An AI Systems Architect’s phased implementation strategy</figcaption></figure><h3>Good News for Backend Engineers: The “90/10 Rule”</h3><p>This spectrum model, especially the functions from the center to the right, is good news for the vast community of web backend engineers. Because for a production-grade AI application, <strong>“90% is the backend engineering we are already familiar with.”</strong></p><p>The overlapping 90% of skills include:</p><ul><li>Infrastructure (Kubernetes, Docker, CI/CD)</li><li>Data Engineering (Databases, Kafka, Redis)</li><li>API &amp; Microservice Development (REST, gRPC)</li><li>Core Software Engineering Practices (Observability, Testing, Design Patterns)</li></ul><p>The critical 10% difference lies in:</p><ul><li><strong>The ability to dance with uncertainty:</strong> Designing systems to manage the probabilistic outputs and potential hallucinations of models.</li><li><strong>AI/LLM “systems intuition”:</strong> Deeply understanding concepts like embeddings, prompt engineering, and context windows, and applying them to system design.</li><li><strong>A data-centric mindset:</strong> Treating data quality and version control as first-class citizens in system design.</li><li><strong>A new paradigm of evaluation and monitoring:</strong> Beyond system metrics, also monitoring model performance, data drift, and API costs.</li></ul><figure><img alt="The 90/10 rule: most of it is familiar backend engineering; the critical difference is the last 10%" src="https://www.imaknas.com/writing/ai-job-spectrum/images/ninety-en.png" width="696" height="269" class="keep-legible"><figcaption>The 90/10 rule: most of it is familiar backend engineering; the critical difference is the last 10%</figcaption></figure><p>Backend engineers are only that final 10% of key knowledge away from becoming the most sought-after AI talent in the market.</p><h3>Conclusion: Finding Your Place on This New Map</h3><p>In the age of AI, career paths are no longer single ladders. We all need to become “strategists” who know how to find our position on the map and plan our path forward.</p><ul><li><strong>For individuals:</strong> This map can help you self-assess. Where are you now? In your future development, do you want to deepen your expertise towards the left of the spectrum, pursuing the pinnacle of technology? Or do you want to expand towards the right, maximizing your business impact?</li><li><strong>For companies:</strong> Companies should use this map to examine their team configurations. Are you only hiring a large number of “AI Engineers” but lack the “Architects” who can define the system blueprint? A clear positioning is what leads to precise hiring.</li></ul><p>On this new continent of AI, we need not only brave explorers but also navigators who can draw the maps and guide the way. No matter where you are on the spectrum, clearly understanding your own position and the value of other roles is the sustainable way to navigate this rapidly evolving market.</p>]]></content:encoded>
  </item>
</channel>
</rss>
