← 回到文章列表Back to writing

,閱讀約 5 分鐘, 5 min read

這篇原文是英文,這裡是中文翻譯。

Nobody made a decision. We just described

What a no-AI spec challenge revealed about how engineers actually think.

Last week I ran a competition with my team.

The rules were simple: one ambiguous requirement, 30 minutes without AI, write a spec. Then hand the spec to Claude Code and let it build. At the end, everyone presents. The team votes on the best output.

I expected the results to tell me who was good at specs. They told me something more revealing than that.

The Setup

The requirement was deliberately vague — modeled closely on work we actually do:

When a data owner publishes a dataset listing and a buyer completes a purchase, the system must deduct payment from the buyer’s wallet, generate a transaction record, update both wallet states, and notify relevant parties. Consider idempotency and failure compensation.

Eight engineers. Individual work. No AI during the spec phase. After that, Claude Code only.

I’d spent the previous two sessions teaching the team about Harness Engineering — the idea that the environment you give an AI matters as much as the prompt. That specs are how you engineer that environment. That the bottleneck isn’t the model.

This was the test.

What I Saw

Some people had nothing to start from.

A few specs were nearly verbatim copies of the requirement. The steps section said: deduct payment, generate transaction record, update wallet states, notify relevant parties. Exactly what the requirement said. No decomposition. No sequencing. No edge cases.

This wasn’t laziness. It was something more fundamental: there was no mental model underneath to translate the requirement into structure. When you take away the tool that would normally help build that structure, there’s nothing left to write.

Deep domain knowledge didn’t help as much as I expected.

One of the most experienced engineers on the team — someone who had spent years working on payment systems — got stuck. Their verbal feedback during the exercise was: the requirement wasn’t defined clearly enough. They needed more context before they could design.

That’s how experienced engineers used to work. Someone writes a PRD. You design against it. If the PRD is unclear, you push back and ask for clarification.

But the challenge wasn’t asking them to design against a complete requirement. It was asking them to resolve the ambiguity themselves — to treat the fuzzy requirement as the problem to be structured, not a blocker to complain about. That’s a different skill. Years of domain knowledge don’t automatically transfer to it.

The winner handed full control to AI and stepped away.

One engineer opened Claude Code in auto mode, let it run, and came back to results. Their spec wasn’t the most precise. But it was structured well enough that the AI produced something that looked complete and convincing. The team voted it the best output.

The delegation itself wasn’t the problem. The problem surfaced at the end, when the output came back.

Most people said the AI added things they hadn’t specified — saga patterns, retry logic, compensating events. And most people admitted they weren’t sure, in the time they had, whether those additions were right or wrong.

That’s the real gap. Not that the spec was incomplete — every spec is incomplete. It’s that when the AI makes a decision on your behalf, you need a prior decision of your own to evaluate it against. If you never decided how failure compensation should work, you have no ground to stand on when the AI proposes a saga. You can accept it or reject it, but either way you’re guessing.

A spec that captures your decisions doesn’t prevent the AI from adding things. It gives you the ability to judge what it adds.

The Pattern Nobody Told Me About

After the spec phase, I asked everyone to share one thing: what was the most important decision you made in your spec?

Nobody answered the question. At first I thought they were avoiding it. Then I realized I had never given them a reason to think in decisions in the first place. The spec template I provided had structure — trigger, steps, edge cases — but no choice points. I asked them to describe a system, then expected them to have made decisions about it. That’s not fair.

Description versus decision

A decision sounds like: “I chose to handle failure at the wallet debit step separately from the credit step, because if debit succeeds and credit fails, the rollback logic is different — and I didn’t want the AI to collapse those into a single failure handler.”

One person got close. Their verbal share included: if debit succeeds but credit fails, retry three times; if retry fails, emit a compensating event and close the transaction. That’s a real decision — they had considered multiple options and chosen one with a reason.

Everyone else described their spec. The description was sometimes accurate. But it wasn’t a decision.

When I reflected on this afterward, I realized the problem wasn’t that they couldn’t articulate decisions. It was that they hadn’t made any. Spec writing, for most of them, meant transcribing their first intuition about how the system should work — not choosing between alternatives.

Why This Matters More Than It Used To

In the old workflow, the absence of explicit decisions was survivable. You wrote code. Someone reviewed it. The decision got surfaced through the review cycle. You had time to course-correct.

AI-assisted development removes that forcing function.

When Claude Code executes your spec in minutes, every assumption you didn’t make explicit becomes a decision the AI made for you. It will fill the gaps confidently. It will produce something that looks complete. And you won’t know which gaps it filled until you’re debugging the result.

The spec isn’t just a planning document anymore. It’s the artifact where your decisions live. If you didn’t make decisions in it, you’re not directing the AI — you’re letting the AI direct itself, and signing off on the output.

What I’m Changing for the Next Session

The exercise revealed a gap I didn’t know how to see before: the difference between describing a system and making decisions about a system.

Next time, the spec template will include explicit choice points. Not just “what are your steps” — but “if debit succeeds and credit fails, which do you choose: retry, rollback, or compensating event? Why?” Forcing them to choose between named alternatives is the only way I know to get people into decision mode rather than description mode.

The voting format is changing too. Peer voting on overall output tends to favor output that looks complete over output that is precise. Next time I’m adding a scoring rubric tied to the three things that actually matter: spec coverage, alignment between spec and output, and skeleton clarity.

And before we start, I’ll spend ten minutes showing them what I found. Not to call anyone out — but because the most clarifying insight from this exercise is also the most useful one:

Most of the team didn’t know they weren’t making decisions. They thought describing the system was the same thing.

That’s the gap AI exposes. Not skill. Not effort. The assumption that your first intuition, written down, is a spec.

I run AI engineering workshops for software teams navigating this transition. If your team is working through the same shift, I’d like to hear about it — cshiauknas@gmail.com

沒有人做決定,我們只是在描述

一場不用 AI 的 spec 挑戰,讓我看見工程師實際上是怎麼思考的。

上週我在團隊裡辦了一場比賽。

規則很簡單:一個模糊的需求,30 分鐘內不准用 AI,寫出一份 spec。接著把 spec 交給 Claude Code,讓它去實作。最後每個人上台報告,由團隊投票選出最好的成果。

我原本以為,結果會告訴我誰比較會寫 spec。沒想到它透露的東西比這更多。

比賽設計

這個需求是我刻意寫得很模糊的,內容幾乎就是照著我們平常實際在做的工作設計:

當資料擁有者上架一份資料集、買家完成購買時,系統必須從買家的錢包扣款、產生一筆交易紀錄、更新雙方的錢包狀態,並通知相關人員。請考慮冪等性與失敗補償。

八位工程師,各自作業。寫 spec 的階段不能用 AI,之後只能用 Claude Code。

在這之前的兩堂課,我都在跟團隊講 Harness Engineering:你給 AI 的環境跟 prompt 一樣重要;spec 就是你打造這個環境的方式;瓶頸不在模型。

這場比賽就是驗收。

我看到了什麼

有些人完全不知道從哪裡開始。

有幾份 spec 幾乎是把需求原封不動抄一遍。步驟那一段寫的是:扣款、產生交易紀錄、更新錢包狀態、通知相關人員。跟需求一字不差。沒有拆解,沒有順序,也沒有邊界情況。

這不是偷懶,而是更根本的問題:底下沒有一個心智模型,能把需求轉成結構。平常幫你搭出結構的工具一拿掉,就沒東西可寫了。

深厚的領域知識,幫助沒有我想像中大。

團隊裡最資深的工程師之一,在支付系統領域做了好幾年,卻卡住了。這位工程師在練習過程中的口頭回饋是:需求定義得不夠清楚,需要更多背景資訊才能開始設計。

資深工程師以前就是這樣工作的。有人寫好 PRD,你照著它設計。PRD 不清楚,就退回去要求對方釐清。

但這次挑戰要的不是照著一份完整的需求去設計,而是要他們自己把模糊的地方釐清:把模糊的需求當成要去結構化的問題,而不是拿來抱怨的障礙。這是另一種能力,多年的領域知識不會自動轉移過來。

贏家把主導權整個交給 AI,然後就走開了。

有位工程師用 auto mode 開了 Claude Code,讓它自己跑,回來直接看結果。這份 spec 不是最精準的,但結構夠好,AI 做出來的東西看起來完整又有說服力。團隊投票選它為最佳成果。

把工作交給 AI 本身不是問題。問題是到最後、產出回來的時候才浮現的。

大部分人都說,AI 加了一些他們沒寫進 spec 的東西:saga pattern、retry 邏輯、補償事件。多數人也承認,在有限的時間裡,他們沒把握這些東西加得對不對。

這才是真正的落差。問題不在 spec 不完整,每份 spec 都不完整。問題在於,當 AI 替你做了決定,你得自己先有一個決定,才有東西可以拿來評估它。如果你從來沒決定過失敗補償該怎麼做,AI 提出 saga 的時候,你就沒有立足點。你可以接受,也可以拒絕,但不管選哪個,你都是在猜。

一份記下你各項決定的 spec,不會阻止 AI 加東西,但它讓你有能力判斷 AI 加的東西對不對。

沒人告訴過我的模式

spec 階段結束後,我請每個人分享一件事:你在 spec 裡做的最重要的決定是什麼?

沒有人回答這個問題。一開始我以為他們在迴避,後來才發現,我從來沒給過他們用「決定」來思考的理由。我提供的 spec 範本有結構(觸發條件、步驟、邊界情況),卻沒有任何選擇點。我要他們描述一個系統,又期待他們已經對這個系統做了決定。這不公平。

描述與決定

一個決定聽起來會像這樣:「我選擇把錢包扣款步驟的失敗,跟入帳步驟的失敗分開處理,因為扣款成功但入帳失敗時,rollback 的邏輯不一樣,我不希望 AI 把它們併成同一個失敗處理器。」

有一個人很接近了。這位同事口頭分享的內容包括:如果扣款成功但入帳失敗,就 retry 三次;retry 還是失敗,就發出補償事件並關閉這筆交易。這是一個真正的決定:對方考慮過幾種做法,然後有理由地選了其中一種。

其他人都在描述自己的 spec。描述有時候很準確,但那不是決定。

事後回想,我發現問題不在他們說不出自己的決定,而是他們根本沒做任何決定。對大多數人來說,寫 spec 就是把自己對系統該怎麼運作的第一直覺抄下來,而不是在幾個選項之間做取捨。

為什麼這件事比以前更重要

在舊的工作流程裡,沒有明確的決定還撐得過去。你寫程式,有人 review,決定會在 review 的過程中浮上檯面,你還有時間修正方向。

AI 輔助開發把這個逼你面對問題的機制拿掉了。

當 Claude Code 幾分鐘內就把你的 spec 執行完,每一個你沒講明的假設,都會變成 AI 替你做的決定。它會很有自信地把空白填滿,做出看起來很完整的東西。而你要到 debug 的時候,才會知道它填了哪些空白。

spec 不再只是一份規劃文件,而是你的決定存放的地方。如果你沒在裡面做決定,你就不是在指揮 AI,而是放任 AI 自己指揮自己,最後再在產出上簽名。

下一堂課我要改什麼

這次練習讓我看見一個以前不知道怎麼看見的落差:描述一個系統,和對一個系統做決定,是兩回事。

下次的 spec 範本會加入明確的選擇點。不只是問「你的步驟是什麼」,而是問「如果扣款成功但入帳失敗,你選哪一個:retry、rollback,還是補償事件?為什麼?」逼他們在具名的選項之間做選擇,是我所知道唯一能讓人從描述模式切換到決策模式的方法。

投票方式也會改。同儕針對整體成果投票,往往會偏好看起來完整的成果,而不是精準的成果。下次我會加上一套評分標準,對準真正重要的三件事:spec 的涵蓋度、spec 與產出的一致性,以及骨架是否清楚。

開始之前,我會先花十分鐘給大家看我的發現。不是要點名誰,而是因為這次練習裡最讓人豁然開朗的洞見,也正是最有用的那一個:

團隊裡大多數人並不知道自己沒在做決定。他們以為描述系統就等於做了決定。

這就是 AI 暴露出來的落差。不是能力,也不是努力,而是以為把第一直覺寫下來就是 spec。

我為正在經歷這波轉型的軟體團隊開設 AI 工程工作坊。如果你的團隊也正面對同樣的轉變,很歡迎來聊聊:cshiauknas@gmail.com

Knas 的文章。想聊聊的話:Written by Knas. To get in touch: Email GitHub LinkedIn