Nobody made a decision. We just described
What a no-AI spec challenge revealed about how engineers actually think.

Last week I ran a competition with my team.
The rules were simple: one ambiguous requirement, 30 minutes without AI, write a spec. Then hand the spec to Claude Code and let it build. At the end, everyone presents. The team votes on the best output.
I expected the results to tell me who was good at specs. They told me something more revealing than that.
The Setup

The requirement was deliberately vague — modeled closely on work we actually do:
When a data owner publishes a dataset listing and a buyer completes a purchase, the system must deduct payment from the buyer’s wallet, generate a transaction record, update both wallet states, and notify relevant parties. Consider idempotency and failure compensation.
Eight engineers. Individual work. No AI during the spec phase. After that, Claude Code only.
I’d spent the previous two sessions teaching the team about Harness Engineering — the idea that the environment you give an AI matters as much as the prompt. That specs are how you engineer that environment. That the bottleneck isn’t the model.
This was the test.
What I Saw
Some people had nothing to start from.
A few specs were nearly verbatim copies of the requirement. The steps section said: deduct payment, generate transaction record, update wallet states, notify relevant parties. Exactly what the requirement said. No decomposition. No sequencing. No edge cases.
This wasn’t laziness. It was something more fundamental: there was no mental model underneath to translate the requirement into structure. When you take away the tool that would normally help build that structure, there’s nothing left to write.
Deep domain knowledge didn’t help as much as I expected.
One of the most experienced engineers on the team — someone who had spent years working on payment systems — got stuck. Their verbal feedback during the exercise was: the requirement wasn’t defined clearly enough. They needed more context before they could design.
That’s how experienced engineers used to work. Someone writes a PRD. You design against it. If the PRD is unclear, you push back and ask for clarification.
But the challenge wasn’t asking them to design against a complete requirement. It was asking them to resolve the ambiguity themselves — to treat the fuzzy requirement as the problem to be structured, not a blocker to complain about. That’s a different skill. Years of domain knowledge don’t automatically transfer to it.
The winner handed full control to AI and stepped away.
One engineer opened Claude Code in auto mode, let it run, and came back to results. Their spec wasn’t the most precise. But it was structured well enough that the AI produced something that looked complete and convincing. The team voted it the best output.
The delegation itself wasn’t the problem. The problem surfaced at the end, when the output came back.
Most people said the AI added things they hadn’t specified — saga patterns, retry logic, compensating events. And most people admitted they weren’t sure, in the time they had, whether those additions were right or wrong.
That’s the real gap. Not that the spec was incomplete — every spec is incomplete. It’s that when the AI makes a decision on your behalf, you need a prior decision of your own to evaluate it against. If you never decided how failure compensation should work, you have no ground to stand on when the AI proposes a saga. You can accept it or reject it, but either way you’re guessing.
A spec that captures your decisions doesn’t prevent the AI from adding things. It gives you the ability to judge what it adds.
The Pattern Nobody Told Me About

After the spec phase, I asked everyone to share one thing: what was the most important decision you made in your spec?
Nobody answered the question. At first I thought they were avoiding it. Then I realized I had never given them a reason to think in decisions in the first place. The spec template I provided had structure — trigger, steps, edge cases — but no choice points. I asked them to describe a system, then expected them to have made decisions about it. That’s not fair.
Description versus decision
A decision sounds like: “I chose to handle failure at the wallet debit step separately from the credit step, because if debit succeeds and credit fails, the rollback logic is different — and I didn’t want the AI to collapse those into a single failure handler.”
One person got close. Their verbal share included: if debit succeeds but credit fails, retry three times; if retry fails, emit a compensating event and close the transaction. That’s a real decision — they had considered multiple options and chosen one with a reason.
Everyone else described their spec. The description was sometimes accurate. But it wasn’t a decision.
When I reflected on this afterward, I realized the problem wasn’t that they couldn’t articulate decisions. It was that they hadn’t made any. Spec writing, for most of them, meant transcribing their first intuition about how the system should work — not choosing between alternatives.
Why This Matters More Than It Used To
In the old workflow, the absence of explicit decisions was survivable. You wrote code. Someone reviewed it. The decision got surfaced through the review cycle. You had time to course-correct.
AI-assisted development removes that forcing function.
When Claude Code executes your spec in minutes, every assumption you didn’t make explicit becomes a decision the AI made for you. It will fill the gaps confidently. It will produce something that looks complete. And you won’t know which gaps it filled until you’re debugging the result.
The spec isn’t just a planning document anymore. It’s the artifact where your decisions live. If you didn’t make decisions in it, you’re not directing the AI — you’re letting the AI direct itself, and signing off on the output.
What I’m Changing for the Next Session
The exercise revealed a gap I didn’t know how to see before: the difference between describing a system and making decisions about a system.
Next time, the spec template will include explicit choice points. Not just “what are your steps” — but “if debit succeeds and credit fails, which do you choose: retry, rollback, or compensating event? Why?” Forcing them to choose between named alternatives is the only way I know to get people into decision mode rather than description mode.
The voting format is changing too. Peer voting on overall output tends to favor output that looks complete over output that is precise. Next time I’m adding a scoring rubric tied to the three things that actually matter: spec coverage, alignment between spec and output, and skeleton clarity.
And before we start, I’ll spend ten minutes showing them what I found. Not to call anyone out — but because the most clarifying insight from this exercise is also the most useful one:
Most of the team didn’t know they weren’t making decisions. They thought describing the system was the same thing.
That’s the gap AI exposes. Not skill. Not effort. The assumption that your first intuition, written down, is a spec.

I run AI engineering workshops for software teams navigating this transition. If your team is working through the same shift, I’d like to hear about it — cshiauknas@gmail.com