AI Raises the Cost of Not Thinking

In March 2026, Andrej Karpathy published a project called autoresearch. The idea: give an AI agent a real LLM training setup and let it experiment overnight. It modifies code, trains for five minutes, checks if the result improved, keeps or discards, and repeats. You wake up to a log of experiments and a better model.
His framing of the human’s role stuck with me: “You’re not touching any of the Python files like you normally would as a researcher. Instead, you are programming the program.”
That’s the shift. Not writing code. Programming the system that writes the code.
I’d been living a version of this for weeks — not in ML research, but in content creation. Every day, an AI agent running on OpenClaw generates scripts, produces videos, publishes them, pulls analytics, and runs A/B experiments on what worked. I’m not in the loop for any of that execution. What I do is define what the system should be optimizing for, catch where it’s going wrong, and redirect it before it drifts.
At some point I realized the pattern underneath all of this had nothing to do with video. It was the same loop Karpathy was describing — just in a different domain. So I extracted it into an open-source framework: adaptive-iteration.
But the more interesting story isn’t what we built. It’s what building it revealed.
The framework wasn’t designed. It emerged.
adaptive-iteration describes a four-part loop: produce something, measure it, generate hypotheses about what to change, then periodically challenge whether you're optimizing for the right thing at all. Repeat.
Here's what's strange: the framework itself was built through exactly that process.
It wasn't designed upfront from first principles. It emerged from running a real system — a content pipeline that had been iterating on itself for weeks — and recognizing a structure that kept reappearing under the domain-specific noise. I described that structure to an AI agent running on OpenClaw. We built the abstraction together. I defined what it should be. The agent implemented it. I caught what was wrong — a dependency direction that was backwards, implementation details that didn't belong in a public repo, architectural decisions that needed to be made explicit. The agent fixed it. That cycle repeated until something clean enough to ship emerged.
The tool was made by the process it describes. That's not a coincidence — it's a signal. The best abstractions don't come from design. They come from running something real long enough to see what's actually there.
Karpathy's autoresearch is a specific instantiation of this loop, hardwired to ML training. adaptive-iteration is the same loop made portable — a Ledger that records what happened, an Analyzer that surfaces what worked, a HypothesisEngine that reasons about what to try next, and a DomainAdapter interface that's the only layer touching your actual system. Bring your own domain. The framework handles the rest.
AI doesn’t lower the cost of building. It raises the cost of not thinking clearly.
Here’s the part the “AI democratizes everything” narrative gets wrong.
Early in the process, I told the agent to push the framework to GitHub. It did. But when I looked at what was actually pushed, one of the adapter files wasn’t a clean example — it was the real implementation, with hardcoded private paths, directly exposing the internals of a production system I hadn’t intended to share. The agent had done exactly what was asked. The problem was what was asked hadn’t been thought through clearly enough.
In the old world, this kind of mistake had friction built in. Building was slow. A fuzzy mental model had time to get clarified during implementation — the cost of the process forced you to think. When the feedback loop is instant, that forcing function disappears. The agent executes your mental model at full speed, including its flaws.
This is the thing nobody says: AI doesn’t make building easier. It makes the quality of your mental model more consequential, faster. A clear model ships something good overnight. A fuzzy model ships something wrong overnight — and it’s already on PyPI before you notice.
The bottleneck didn’t disappear. It transformed. And it’s less forgiving than the old one.
The governance asymmetry
Which brings us to the part of the paradigm shift that’s less comfortable to say out loud.
“Anyone can build now” is technically true. But what most of these conversations miss is that building and building the right thing are different problems. The first is an execution problem. The second is a thinking problem. AI solves the first. It amplifies the second.
The new bottleneck — recognizing generalizable patterns across domains, holding a clear mental model of what a system should be, knowing which architectural decisions matter and which don’t — isn’t more democratically distributed than the ability to code. It’s just different. A person who can recognize that their content optimization loop and their proposal strategy loop share the same underlying structure, and can govern an AI agent to extract that abstraction cleanly, will compound enormously from these tools. Someone who can’t will build faster and make more mistakes faster.
What’s genuinely democratized is the ceiling. A researcher who has spent years in a domain but never learned to code can now build the analysis tools their field has needed. A teacher can build adaptive curriculum systems. A founder can build the product they’ve been describing to developers who never quite got it right. The ideas that stayed ideas because the person who had them couldn’t implement them — those ideas are now executable.
But the floor didn’t move. Governance — the ability to define what should be built, evaluate what was built, and catch the gap between the two — is still a skill. It still has to be developed. Karpathy’s autoresearch works because Karpathy is Karpathy. The program.md he writes to direct his agents carries decades of ML research intuition. The output quality is a function of the governance quality.
That’s the paradigm shift I actually care about. Not that execution is cheap. That governance is now the core competency — and unlike code, nobody’s teaching it yet.
Three questions I deliberately left open in this piece.
What happens in domains where “Measure” is the hard problem — where feedback cycles are long, noisy, or fundamentally contested? The framework assumes you can observe outcomes. A lot of the highest-value decisions in the world don’t come with a loss function.
What does governance actually decompose into as a skill? I called it a competency and then left it there. That’s not good enough. The difference between intuition and a teachable framework is whether you can name the components.
And the one that keeps me up at night: AI is already pushing into the governance layer — LLM-as-a-Judge, agentic workflows with self-correction, systems that generate their own hypotheses. The boundary I described between human governance and AI execution is already moving. What do you govern when the thing you’re governing is itself governing something else?
I built a framework that claims to be domain-agnostic. I’m not sure I’ve tested that claim honestly. I’m still working on all three.
The repo
adaptive-iteration is a domain-agnostic experimentation framework. Bring your own domain — the framework handles the experiment → measure → learn → challenge loop.
→ github.com/imaknas/adaptive-iteration
→ pypi.org/project/adaptive-iteration