← 回到文章列表Back to writing

,閱讀約 5 分鐘, 5 min read

adaptive-iteration 後來改名為 Ordal,這篇寫於改名之前。adaptive-iteration has since been renamed Ordal; this piece predates the rename.

這篇原文是英文,這裡是中文翻譯。

AI Raises the Cost of Not Thinking

In March 2026, Andrej Karpathy published a project called autoresearch. The idea: give an AI agent a real LLM training setup and let it experiment overnight. It modifies code, trains for five minutes, checks if the result improved, keeps or discards, and repeats. You wake up to a log of experiments and a better model.

His framing of the human’s role stuck with me: “You’re not touching any of the Python files like you normally would as a researcher. Instead, you are programming the program.”

That’s the shift. Not writing code. Programming the system that writes the code.

I’d been living a version of this for weeks — not in ML research, but in content creation. Every day, an AI agent running on OpenClaw generates scripts, produces videos, publishes them, pulls analytics, and runs A/B experiments on what worked. I’m not in the loop for any of that execution. What I do is define what the system should be optimizing for, catch where it’s going wrong, and redirect it before it drifts.

At some point I realized the pattern underneath all of this had nothing to do with video. It was the same loop Karpathy was describing — just in a different domain. So I extracted it into an open-source framework: adaptive-iteration.

But the more interesting story isn’t what we built. It’s what building it revealed.

The framework wasn’t designed. It emerged.

adaptive-iteration describes a four-part loop: produce something, measure it, generate hypotheses about what to change, then periodically challenge whether you're optimizing for the right thing at all. Repeat.

Here's what's strange: the framework itself was built through exactly that process.

It wasn't designed upfront from first principles. It emerged from running a real system — a content pipeline that had been iterating on itself for weeks — and recognizing a structure that kept reappearing under the domain-specific noise. I described that structure to an AI agent running on OpenClaw. We built the abstraction together. I defined what it should be. The agent implemented it. I caught what was wrong — a dependency direction that was backwards, implementation details that didn't belong in a public repo, architectural decisions that needed to be made explicit. The agent fixed it. That cycle repeated until something clean enough to ship emerged.

The tool was made by the process it describes. That's not a coincidence — it's a signal. The best abstractions don't come from design. They come from running something real long enough to see what's actually there.

Karpathy's autoresearch is a specific instantiation of this loop, hardwired to ML training. adaptive-iteration is the same loop made portable — a Ledger that records what happened, an Analyzer that surfaces what worked, a HypothesisEngine that reasons about what to try next, and a DomainAdapter interface that's the only layer touching your actual system. Bring your own domain. The framework handles the rest.

AI doesn’t lower the cost of building. It raises the cost of not thinking clearly.

Here’s the part the “AI democratizes everything” narrative gets wrong.

Early in the process, I told the agent to push the framework to GitHub. It did. But when I looked at what was actually pushed, one of the adapter files wasn’t a clean example — it was the real implementation, with hardcoded private paths, directly exposing the internals of a production system I hadn’t intended to share. The agent had done exactly what was asked. The problem was what was asked hadn’t been thought through clearly enough.

In the old world, this kind of mistake had friction built in. Building was slow. A fuzzy mental model had time to get clarified during implementation — the cost of the process forced you to think. When the feedback loop is instant, that forcing function disappears. The agent executes your mental model at full speed, including its flaws.

This is the thing nobody says: AI doesn’t make building easier. It makes the quality of your mental model more consequential, faster. A clear model ships something good overnight. A fuzzy model ships something wrong overnight — and it’s already on PyPI before you notice.

The bottleneck didn’t disappear. It transformed. And it’s less forgiving than the old one.

The governance asymmetry

Which brings us to the part of the paradigm shift that’s less comfortable to say out loud.

“Anyone can build now” is technically true. But what most of these conversations miss is that building and building the right thing are different problems. The first is an execution problem. The second is a thinking problem. AI solves the first. It amplifies the second.

The new bottleneck — recognizing generalizable patterns across domains, holding a clear mental model of what a system should be, knowing which architectural decisions matter and which don’t — isn’t more democratically distributed than the ability to code. It’s just different. A person who can recognize that their content optimization loop and their proposal strategy loop share the same underlying structure, and can govern an AI agent to extract that abstraction cleanly, will compound enormously from these tools. Someone who can’t will build faster and make more mistakes faster.

What’s genuinely democratized is the ceiling. A researcher who has spent years in a domain but never learned to code can now build the analysis tools their field has needed. A teacher can build adaptive curriculum systems. A founder can build the product they’ve been describing to developers who never quite got it right. The ideas that stayed ideas because the person who had them couldn’t implement them — those ideas are now executable.

But the floor didn’t move. Governance — the ability to define what should be built, evaluate what was built, and catch the gap between the two — is still a skill. It still has to be developed. Karpathy’s autoresearch works because Karpathy is Karpathy. The program.md he writes to direct his agents carries decades of ML research intuition. The output quality is a function of the governance quality.

That’s the paradigm shift I actually care about. Not that execution is cheap. That governance is now the core competency — and unlike code, nobody’s teaching it yet.

Three questions I deliberately left open in this piece.

What happens in domains where “Measure” is the hard problem — where feedback cycles are long, noisy, or fundamentally contested? The framework assumes you can observe outcomes. A lot of the highest-value decisions in the world don’t come with a loss function.

What does governance actually decompose into as a skill? I called it a competency and then left it there. That’s not good enough. The difference between intuition and a teachable framework is whether you can name the components.

And the one that keeps me up at night: AI is already pushing into the governance layer — LLM-as-a-Judge, agentic workflows with self-correction, systems that generate their own hypotheses. The boundary I described between human governance and AI execution is already moving. What do you govern when the thing you’re governing is itself governing something else?

I built a framework that claims to be domain-agnostic. I’m not sure I’ve tested that claim honestly. I’m still working on all three.

The repo

adaptive-iteration is a domain-agnostic experimentation framework. Bring your own domain — the framework handles the experiment → measure → learn → challenge loop.

→ github.com/imaknas/adaptive-iteration
→ pypi.org/project/adaptive-iteration

AI 讓不思考的代價變高了

2026 年 3 月,Andrej Karpathy 發表了一個叫 autoresearch 的專案。概念是:給 AI agent 一套真實的 LLM 訓練環境,讓它整晚自己做實驗。它會修改程式碼、訓練五分鐘、檢查結果有沒有變好、決定保留或捨棄,然後重複。你早上醒來,就會拿到一份實驗紀錄和一個更好的模型。

他對人類角色的描述讓我印象很深:「你不會像一般研究者那樣去動任何 Python 檔案。你是在為那個程式寫程式。」

這就是轉變所在。不是寫程式碼,而是為寫程式碼的系統寫程式。

這種生活的某個版本,我已經過了好幾週,只不過不是在 ML 研究,而是在內容創作。每天,一個跑在 OpenClaw 上的 AI agent 會產生腳本、製作影片、發布上線、抓取數據,再針對哪些內容有效做 A/B 實驗。這些執行環節我都不參與。我做的是定義系統應該優化什麼、抓出它哪裡出錯,並在它偏掉之前把它導回來。

某個時間點我發現,這一切底下的模式其實跟影片毫無關係。它就是 Karpathy 描述的那個迴圈,只是換了一個領域。所以我把它抽出來,做成一個開源框架:adaptive-iteration。

但比起我們做了什麼,更有意思的是做的過程揭露了什麼。

這個框架不是設計出來的,而是自己長出來的。

adaptive-iteration 描述的是一個四段式迴圈:產出某個東西、衡量它、針對要改什麼提出假設,然後定期質疑你到底有沒有在優化對的東西。然後重複。

奇妙的是:這個框架本身,就是透過這個過程做出來的。

它不是一開始就從第一原理設計好的。它來自一個真實運作中的系統(一條已經自我迭代了好幾週的內容產線),我在一堆領域特有的雜訊底下,認出一個反覆出現的結構。我把這個結構描述給一個跑在 OpenClaw 上的 AI agent,我們一起把抽象做出來。我定義它應該是什麼,agent 負責實作。我抓出哪裡不對:一個方向反了的依賴關係、不該出現在公開 repo 裡的實作細節、需要明確寫出來的架構決策。agent 再去修正。這個循環一再重複,直到做出一個夠乾淨、可以發布的東西。

這個工具,是由它所描述的過程做出來的。這不是巧合,而是一個訊號。最好的抽象不是設計出來的,而是把一個真實的東西跑得夠久,久到看清它真正的樣子。

Karpathy 的 autoresearch 是這個迴圈的一個具體實例,綁死在 ML 訓練上。adaptive-iteration 則是把同一個迴圈做成可以帶著走的版本:記錄發生了什麼的 Ledger、找出什麼有效的 Analyzer、推理下一步該試什麼的 HypothesisEngine,以及 DomainAdapter 介面,也就是唯一會碰到你實際系統的那一層。領域你自己帶,其餘交給框架。

AI 沒有降低打造東西的成本,它提高的是想不清楚的代價。

這正是「AI 讓一切民主化」這套說法搞錯的地方。

在過程初期,我叫 agent 把框架推上 GitHub,它照做了。但我去看實際推上去的內容時,發現其中一個 adapter 檔案不是乾淨的範例,而是真正的實作,裡面寫死了私人路徑,直接暴露出一個我原本沒打算公開的正式環境系統的內部細節。agent 完全照著指示做,問題出在指示本身沒有想清楚。

在過去,這類錯誤本身就帶有摩擦力。打造東西很慢,模糊的心智模型有時間在實作過程中被釐清:過程的成本逼著你思考。當回饋迴圈變成即時的,這個逼你思考的機制就消失了。agent 會全速執行你的心智模型,連同它的缺陷一起。

這是沒人講的事:AI 沒有讓打造東西變容易,而是讓你心智模型的品質影響更大,後果也來得更快。清楚的模型一個晚上就能交出好東西;模糊的模型一個晚上就能交出錯的東西,而且在你發現之前,它已經上了 PyPI。

瓶頸沒有消失,只是換了形式,而且比舊的更不留情面。

治理的不對稱

這就帶到了這場典範轉移裡,比較不好大聲說出口的部分。

「現在人人都能打造東西」,技術上沒錯。但大多數這類討論忽略了一點:打造東西,跟打造對的東西,是兩個不同的問題。前者是執行問題,後者是思考問題。AI 解決了前者,卻放大了後者。

新的瓶頸(辨識跨領域可通用的模式、對一個系統該是什麼樣子保有清楚的心智模型、知道哪些架構決策重要而哪些不重要)並沒有比寫程式的能力分布得更平均,只是不一樣而已。一個人如果能看出自己的內容優化迴圈和提案策略迴圈共享同一個底層結構,又能駕馭 AI agent 把這個抽象乾淨地抽出來,這樣的人從這些工具得到的回報會不斷複利放大。做不到的人,只會做得更快,也錯得更快。

真正被民主化的是天花板。一位在某個領域深耕多年、卻從沒學過寫程式的研究者,現在可以做出自己領域一直需要的分析工具。老師可以打造適性化的課程系統。創辦人可以親手做出那個自己一直跟開發者描述、卻始終沒被做對的產品。那些因為提出的人無法實作、只能停留在想法階段的點子,現在都能執行了。

但地板沒有動。治理,也就是定義該打造什麼、評估打造出來的東西、並抓出兩者之間落差的能力,依然是一種技能,依然需要培養。Karpathy 的 autoresearch 之所以行得通,是因為 Karpathy 就是 Karpathy。他寫來指揮 agent 的 program.md,承載了數十年的 ML 研究直覺。產出的品質,取決於治理的品質。

這才是我真正在意的典範轉移。不是執行變便宜了,而是治理如今成了核心能力,而且跟寫程式不同,目前還沒有人在教。

這篇文章裡,我刻意留下三個沒有答案的問題。

在「衡量」本身就是難題的領域,也就是回饋週期很長、雜訊很多,或本質上就充滿爭議的情況下,會發生什麼事?這個框架假設你能觀察到結果,但世界上很多價值最高的決策,並不會附上一個 loss function。

治理作為一種技能,實際上可以拆解成哪些部分?我稱它為一種核心能力,然後就停在那裡了。這樣不夠。直覺和可以教的框架之間的差別,就在於你能不能說出它由哪些部分組成。

還有一個讓我睡不著的問題:AI 已經在往治理層推進了,像是 LLM-as-a-Judge、會自我修正的 agentic workflow、會自己產生假設的系統。我描述的那條人類治理與 AI 執行之間的界線,已經在移動了。當你治理的對象本身也在治理別的東西,你治理的究竟是什麼?

我做了一個號稱不限領域的框架,但我不確定自己有沒有誠實地檢驗過這個說法。這三個問題,我都還在想。

Repo

adaptive-iteration 是一個不限領域的實驗框架。領域你自己帶,實驗 → 衡量 → 學習 → 質疑的迴圈交給框架處理。

→ github.com/imaknas/adaptive-iteration
→ pypi.org/project/adaptive-iteration

Knas 的文章。想聊聊的話:Written by Knas. To get in touch: Email GitHub LinkedIn