← 回到文章列表Back to writing

,閱讀約 3 分鐘, 3 min read

這篇原文是英文,這裡是中文翻譯。

Why Your Team Is Using AI Wrong — And How One Session Changed That

The missing skill isn’t prompting. It’s Harness Engineering.

Last week I ran an AI coding session at my company, walking the team through how I actually use Claude Code day to day.

What I discovered surprised me: the gap between people getting 10x productivity and people feeling like AI is overhyped isn’t about skill, intelligence, or even experience with AI tools. It’s about exposure to a good working model.

Most people have never seen what high-efficiency AI-assisted coding looks like. So they don’t know what they’re missing.

The concept: Harness Engineering

The industry is converging on a term for this: harness engineering. Both Anthropic and OpenAI engineering teams have published on the practice — the core idea being that engineering the environment an agent operates in matters as much as the prompts you give it. The insight is simple but profound: a capable model running in a poorly designed environment will consistently underperform. The harness — the initialization, the constraints, the structured context — determines whether your agent stays on task or goes off the rails. Most developers skip this entirely. They open a chat window and start prompting.

The workflow I showed them has three steps.

Step 1: /init — build the harness first

Before writing a single line of code, run /init in Claude Code. This is your harness initialization: the agent reads your project structure, tech constraints, and coding conventions, and creates a shared context that every subsequent session inherits. This is the difference between an agent that understands your codebase and one that's making educated guesses. Every engineer on the team working with the same initialized harness gets consistent, codebase-aware output. Skip this, and you're not doing AI-assisted development — you're doing expensive autocomplete.

Step 2: ticket → spec — clarify before you execute

A Jira ticket tells you what to build. A spec tells the agent how to know when it’s done. Before any implementation, work with the agent to convert the ticket into a proper spec: acceptance criteria, edge cases, integration points. During this process, the agent surfaces gaps in the requirements — things you hadn’t fully thought through. This isn’t extra work. It’s the planning that should have happened anyway, now with an agent that actively helps you find the holes. Anthropic’s research found that decomposing work into clearly-scoped units — rather than one-shotting a full feature — was one of the biggest levers for reliable agent performance. The spec is how you create those units.

Step 3: review the plan, then let go

Once the spec is solid, the agent generates an implementation plan. Review the plan — not the code, just the direction. Is the approach right? Does it account for what you care about? When you’re satisfied, let it run. Human sets direction. Agent handles execution. This maps directly to what Anthropic calls the “coding agent” pattern: give the agent a clear scope, let it make incremental progress, check back when it’s done.

The real insight

AI coding tools don’t underperform because the AI isn’t good enough. They underperform because we hand them vague inputs and expect precise outputs. The engineers on my team who already had strong spec discipline picked this up immediately. Those used to “figuring it out as they go” found it harder — not because of the AI, but because it exposed gaps in how explicitly they were thinking about requirements. That’s the uncomfortable truth about AI productivity: it amplifies your existing engineering habits. Good habits scale. Vague habits stay vague, just faster.

If you’re leading a team using AI coding tools

Don’t just give people access and hope for the best. Run a session. Show them what a properly initialized harness looks like. Let them see the difference between dropping a ticket into Claude and building a spec together first. The bottleneck isn’t the model. It’s whether your team knows how to engineer the environment it runs in.

你的團隊為什麼用錯了 AI:一堂課帶來的改變

缺的不是下 prompt 的技巧,而是 Harness Engineering。

上週我在公司帶了一堂 AI coding 課,讓團隊看看我平常實際上是怎麼用 Claude Code 的。

我的發現讓我很意外:有些人用 AI 拿到 10x 的生產力,有些人覺得 AI 被過度炒作,兩者之間的差距跟技術能力、聰明程度,甚至跟使用 AI 工具的經驗都無關,而是在於有沒有看過一個好的工作模式。

大多數人從來沒看過高效率的 AI 輔助開發長什麼樣子,所以也不知道自己錯過了什麼。

核心概念:Harness Engineering

業界對這件事漸漸有了共同的名稱:harness engineering。Anthropic 和 OpenAI 的工程團隊都發表過相關的做法,核心概念是:為 agent 打造它運作的環境,跟你給它的 prompt 一樣重要。這個洞見簡單卻深刻:一個能力很強的模型,放在設計不良的環境裡,表現就是會一直不如預期。harness(初始化、限制條件、結構化的脈絡)決定了你的 agent 是能專注在任務上,還是會整個跑偏。大多數開發者完全跳過這一步,打開聊天視窗就開始下 prompt。

我示範的工作流程有三個步驟。

步驟 1:/init,先把 harness 建好

在寫任何一行程式碼之前,先在 Claude Code 裡跑 /init。這就是 harness 的初始化:agent 會讀取你的專案結構、技術限制和程式碼慣例,建立一份共享的脈絡,之後的每個 session 都會繼承它。有沒有這一步,差別就在於 agent 是真的理解你的程式碼庫,還是只能憑經驗猜。團隊裡每個工程師都用同一套初始化好的 harness,就能得到一致、了解程式碼庫的產出。跳過這一步,你做的就不是 AI 輔助開發,而是很昂貴的自動補全。

步驟 2:ticket → spec,先釐清再執行

Jira ticket 告訴你要做什麼,spec 則告訴 agent 怎樣才算做完。在開始實作之前,先跟 agent 一起把 ticket 轉成一份像樣的 spec:驗收條件、邊界情況、整合點。在這個過程中,agent 會把需求裡的缺口挖出來,也就是那些你還沒完全想清楚的地方。這不是額外的工作,而是本來就該做的規劃,只是現在有個 agent 主動幫你找漏洞。Anthropic 的研究發現,把工作拆成範圍明確的單位,而不是一次把整個功能生出來,是讓 agent 表現穩定的最大槓桿之一。spec 就是你切出這些單位的方式。

步驟 3:審核計畫,然後放手

spec 確定之後,agent 會產生一份實作計畫。審核這份計畫,不看程式碼,只看方向。做法對嗎?有沒有顧到你在意的事?滿意了就讓它跑。人決定方向,agent 負責執行。這正好對應到 Anthropic 所說的「coding agent」模式:給 agent 明確的範圍,讓它一步步推進,做完再回來檢查。

真正的洞見

AI coding 工具表現不好,不是因為 AI 不夠強,而是因為我們丟給它模糊的輸入,卻期待精準的輸出。團隊裡原本就有紮實 spec 紀律的工程師,馬上就上手了。習慣「邊做邊想」的人則覺得比較難,原因不在 AI,而是它暴露出他們思考需求時有多不明確。這就是 AI 生產力讓人不太舒服的真相:它會放大你既有的工程習慣。好習慣會跟著規模化,模糊的習慣還是一樣模糊,只是變快了。

如果你正在帶一個使用 AI coding 工具的團隊

不要只是開權限給大家,然後祈禱一切順利。辦一堂課,讓他們看看一個正確初始化的 harness 長什麼樣子,讓他們親眼看到,把 ticket 直接丟給 Claude,跟先一起把 spec 寫好,差別在哪裡。瓶頸不在模型,而在你的團隊知不知道怎麼打造它運作的環境。

Knas 的文章。想聊聊的話:Written by Knas. To get in touch: Email GitHub LinkedIn