[{"data":1,"prerenderedAt":792},["ShallowReactive",2],{"blog-spec-driven-development-coding-agents-en":3},{"slug":4,"published":5,"minutes":6,"category":7,"tags":8,"keywords":14,"about":23,"sources":35,"cover":78,"og":79,"expertise":80,"locales":81,"lang":82,"title":85,"description":86,"coverAlt":87,"metaTitle":88,"takeaways":89,"faq":96,"toc":109,"blocks":140,"others":517},"spec-driven-development-coding-agents","2026-09-30",11,"agents",[9,10,11,12,13],"Spec-driven development","Coding agents","Acceptance criteria","Plan mode","AI engineering",[15,16,17,18,19,20,21,22],"spec-driven development","spec driven development coding agents","acceptance criteria for coding agents","AI coding agent plan mode","GitHub Spec Kit","Kiro specs requirements design tasks","vibe coding vs spec-driven development","EARS requirements syntax",[24,27,29,32],{"name":25,"url":26},"Vibe coding","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FVibe_coding",{"name":19,"url":28},"https:\u002F\u002Fgithub.com\u002Fgithub\u002Fspec-kit",{"name":30,"url":31},"EARS (Easy Approach to Requirements Syntax)","https:\u002F\u002Falistairmavin.com\u002Fears\u002F",{"name":33,"url":34},"Kiro","https:\u002F\u002Fkiro.dev",[36,39,41,44,47,50,52,55,58,61,64,67,70,72,75],{"title":37,"url":38},"Birgitta Böckeler: Understanding Spec-Driven-Development: Kiro, spec-kit, and Tessl (martinfowler.com, 15 October 2025)","https:\u002F\u002Fmartinfowler.com\u002Farticles\u002Fexploring-gen-ai\u002Fsdd-3-tools.html",{"title":40,"url":28},"GitHub Spec Kit: README",{"title":42,"url":43},"GitHub Spec Kit: spec-driven.md methodology","https:\u002F\u002Fgithub.com\u002Fgithub\u002Fspec-kit\u002Fblob\u002Fmain\u002Fspec-driven.md",{"title":45,"url":46},"Kiro documentation: specs","https:\u002F\u002Fkiro.dev\u002Fdocs\u002Fspecs\u002F",{"title":48,"url":49},"Kiro: Introducing Kiro","https:\u002F\u002Fkiro.dev\u002Fblog\u002Fintroducing-kiro\u002F",{"title":51,"url":31},"Alistair Mavin: EARS, the Easy Approach to Requirements Syntax",{"title":53,"url":54},"Anthropic: Best practices for Claude Code","https:\u002F\u002Fcode.claude.com\u002Fdocs\u002Fen\u002Fbest-practices",{"title":56,"url":57},"Cline documentation: Plan and Act","https:\u002F\u002Fdocs.cline.bot\u002Ffeatures\u002Fplan-and-act",{"title":59,"url":60},"OpenAI: Codex slash command reference","https:\u002F\u002Flearn.chatgpt.com\u002Fdocs\u002Freference\u002Fslash-commands",{"title":62,"url":63},"METR: early-2025 AI and experienced open-source developer productivity (July 2025)","https:\u002F\u002Fmetr.org\u002Fblog\u002F2025-07-10-early-2025-ai-experienced-os-dev-study\u002F",{"title":65,"url":66},"METR: We are changing our developer productivity experiment design (February 2026)","https:\u002F\u002Fmetr.org\u002Fblog\u002F2026-02-24-uplift-update\u002F",{"title":68,"url":69},"Thoughtworks Technology Radar: Spec-driven development","https:\u002F\u002Fwww.thoughtworks.com\u002Fradar\u002Ftechniques\u002Fspec-driven-development",{"title":71,"url":26},"Wikipedia: Vibe coding",{"title":73,"url":74},"Anthropic: Claude API pricing","https:\u002F\u002Fplatform.claude.com\u002Fdocs\u002Fen\u002Fabout-claude\u002Fpricing",{"title":76,"url":77},"Anthropic: Prompt caching","https:\u002F\u002Fplatform.claude.com\u002Fdocs\u002Fen\u002Fbuild-with-claude\u002Fprompt-caching","\u002Fimages\u002Fblog\u002Fspec-driven-development-coding-agents\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fspec-driven-development-coding-agents\u002Fog.jpg","ai-engineer",[82,83,84],"en","de","hu","Spec-driven development for coding agents: agree the plan before the code","Vibe coding breaks on real codebases. Write a spec with acceptance criteria, a plan and tasks, let the agent tick them off, and review before the first line of code.","Cover art for spec-driven development: a pipeline from spec and plan to tasks and verification, with a review gate before any code is written.","Spec-driven development for coding agents · Balázs Csorba",[90,91,92,93,94,95],"Vibe coding breaks on real codebases because the rules nobody wrote down are the ones the agent gets wrong. Write the intent down before the agent writes code.","A useful spec states behaviour as testable acceptance criteria, names what is out of scope and ends with a check that proves the feature works.","Approve the plan before any edit: the files, the interfaces, the order of work and the risks. It is the cheapest place to catch a wrong design.","Read each tool for what it keeps. Spec Kit and Kiro keep spec files, while plan modes in Claude Code and Cline give a read-only planning step, not a durable spec.","Spec tokens are cheap, but review time and attention are not. Match the process to the change, and skip the plan for a diff you can describe in one sentence.","Judge the result by rework and review time, not by how fast the first draft appeared. METR's measurements show how unreliable that feeling is.",[97,100,103,106],{"q":98,"a":99},"What is spec-driven development?","It means writing down what a feature must do, how it will be built and how you will check it, and only then letting a coding agent write code against those documents. The spec, the plan and the task list are files you can review, diff and test, instead of a chat history.",{"q":101,"a":102},"Is spec-driven development just vibe coding with more paperwork?","More paperwork, yes, but the review moves. Vibe coding can accept output without thorough review. Spec-driven work puts the review on the spec, the plan and the evidence. It costs more for the same feature, so it pays off for changes that span several files or touch risky paths, and not for small fixes.",{"q":104,"a":105},"Which tool should I use?","The one your team will keep using. Spec Kit runs the spec, plan, tasks and implement steps as chat commands and keeps the artefacts in a folder per feature. Kiro keeps requirements, design and tasks as files for each spec. Plan modes in Claude Code and Cline give you a read-only planning step, which helps, but they do not keep a spec for you.",{"q":107,"a":108},"How do acceptance criteria help with code review?","Each criterion becomes a check the reviewer can look at. The reviewer asks whether every criterion has evidence, whether a check can fail, and whether anything outside the scope changed. That is a smaller and more reliable job than reading a large diff without knowing what it was meant to do.",[110,113,116,119,122,125,128,131,134,137],{"id":111,"title":112},"why-vibe-coding-breaks","Why vibe coding breaks on real codebases",{"id":114,"title":115},"the-workflow","The workflow: spec, plan, tasks, verification",{"id":117,"title":118},"spec-template","A spec with acceptance criteria",{"id":120,"title":121},"plan-and-tasks","The plan file and the task checklist",{"id":123,"title":124},"tooling","What the tools really give you",{"id":126,"title":127},"reviewable-and-testable","How specs make agent output reviewable and testable",{"id":129,"title":130},"cost-and-time","Cost and time: when the spec pays off",{"id":132,"title":133},"where-it-falls-short","Where spec-driven development falls short",{"id":135,"title":136},"first-steps","What I would do first",{"id":138,"title":139},"sources","Sources",[141,145,148,151,154,179,182,189,190,193,202,229,232,233,236,239,242,243,246,248,257,259,260,263,319,322,325,327,330,331,334,341,344,350,358,361,362,365,368,371,415,418,419,436,439,440,453,468,469],{"type":142,"content":143},"paragraph",[144],"Vibe coding works until a codebase has history. A prototype can be judged by whether it runs. A production service cannot, because the agent does not know the rules nobody wrote down: the retry rule the payments team agreed on, the flag that protects an old import, the query that must stay under a timeout. It fills those gaps with plausible guesses, and you find them in review, after the code exists. The fix is to write the intent down first: a spec with testable acceptance criteria, a plan you approve, and a task list the agent works through and ticks off, with evidence for each tick. That is spec-driven development, and I would use it for any change that touches more than one module.",{"type":142,"content":146},[147],"Andrej Karpathy coined vibe coding in February 2025 for building software by describing it to a model and accepting the output without thorough review. For a weekend tool that trade is fine. It gets expensive when the habit meets a codebase with conventions, shared modules and paying customers, because then review does the real work, and nobody has defined what it should check.",{"type":149,"level":150,"id":111,"text":112},"heading",2,{"type":142,"content":152},[153],"Four failure patterns keep coming back. They share one cause: the agent works from a prompt, while the team works from an understanding that was never written down.",{"type":155,"ordered":156,"items":157},"list",false,[158,164,169,174],[159,163],{"tag":160,"children":161},"strong",[162],"Invented requirements."," The agent fills gaps with plausible behaviour, such as a new default or an error message nobody agreed on, and nobody decided it.",[165,168],{"tag":160,"children":166},[167],"Convention drift."," Each change is correct in isolation and ignores the patterns the rest of the code relies on.",[170,173],{"tag":160,"children":171},[172],"Undefined done."," Anthropic's guidance for Claude Code notes that without a check the agent can run, “looks done” is the only signal available.",[175,178],{"tag":160,"children":176},[177],"Expensive review."," A large diff written in one pass forces the reviewer to rebuild the intent from the code.",{"type":142,"content":180},[181],"The measured cost is not what most of us expect. In July 2025, METR randomly assigned 246 real issues from large open-source repositories to be done with or without AI tools, mostly Cursor Pro with Claude 3.5 or 3.7 Sonnet, across 16 experienced developers. Before the study they expected AI to make them 24% faster. Afterwards they still believed it had made them 20% faster. The measured result was 19% longer task times with AI allowed.",{"type":183,"variant":184,"title":185,"body":186},"callout","note","Read the 2025 number next to METR's update",[187],[188],"METR says its results are out of date and points to a follow-up from February 2026, with 57 developers and more than 800 tasks. Its estimates there come with confidence intervals that include zero, and METR calls the new data “an unreliable signal”. The lesson for me is about measurement: felt speed is a poor instrument.",{"type":149,"level":150,"id":114,"text":115},{"type":142,"content":191},[192],"The workflow has five stages, and each one produces an artefact a person can read before the next stage starts. The paperwork is not the point. Each gate is a cheap place to stop the agent, and everything between two gates is the agent's job.",{"type":194,"attrs":195,"inner":199,"caption":200},"diagram",{"viewBox":196,"role":197,"aria-labelledby":198},"0 0 720 270","img","d1-sdd-t d1-sdd-d","\u003Ctitle id=\"d1-sdd-t\">Spec-driven workflow with three human gates\u003C\u002Ftitle>\u003Cdesc id=\"d1-sdd-d\">Five boxes in a row from left to right: spec, plan, tasks, implementation and verification. Human gates sit under the spec, the plan and the verification box. Each box names the artefact it produces or the work it does.\u003C\u002Fdesc>\u003Ctext x=\"20\" y=\"24\" class=\"d-label\">Five stages, three gates\u003C\u002Ftext>\u003Crect x=\"20\" y=\"70\" width=\"120\" height=\"84\" rx=\"10\" class=\"d-gold\" \u002F>\u003Ctext x=\"80\" y=\"104\" text-anchor=\"middle\" class=\"d-text\">Spec\u003C\u002Ftext>\u003Ctext x=\"80\" y=\"126\" text-anchor=\"middle\" class=\"d-small\">what, why\u003C\u002Ftext>\u003Crect x=\"160\" y=\"70\" width=\"120\" height=\"84\" rx=\"10\" class=\"d-gold\" \u002F>\u003Ctext x=\"220\" y=\"104\" text-anchor=\"middle\" class=\"d-text\">Plan\u003C\u002Ftext>\u003Ctext x=\"220\" y=\"126\" text-anchor=\"middle\" class=\"d-small\">files and risks\u003C\u002Ftext>\u003Crect x=\"300\" y=\"70\" width=\"120\" height=\"84\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"360\" y=\"104\" text-anchor=\"middle\" class=\"d-text\">Tasks\u003C\u002Ftext>\u003Ctext x=\"360\" y=\"126\" text-anchor=\"middle\" class=\"d-small\">one check each\u003C\u002Ftext>\u003Crect x=\"440\" y=\"70\" width=\"120\" height=\"84\" rx=\"10\" class=\"d-sky\" \u002F>\u003Ctext x=\"500\" y=\"104\" text-anchor=\"middle\" class=\"d-text\">Implement\u003C\u002Ftext>\u003Ctext x=\"500\" y=\"126\" text-anchor=\"middle\" class=\"d-small\">agent ticks off\u003C\u002Ftext>\u003Crect x=\"580\" y=\"70\" width=\"120\" height=\"84\" rx=\"10\" class=\"d-mint\" \u002F>\u003Ctext x=\"640\" y=\"104\" text-anchor=\"middle\" class=\"d-text\">Verify\u003C\u002Ftext>\u003Ctext x=\"640\" y=\"126\" text-anchor=\"middle\" class=\"d-small\">evidence, review\u003C\u002Ftext>\u003Cpath d=\"M140 112 H151\" class=\"d-line\" \u002F>\u003Cpath d=\"M160 112 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Cpath d=\"M280 112 H291\" class=\"d-line\" \u002F>\u003Cpath d=\"M300 112 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Cpath d=\"M420 112 H431\" class=\"d-line\" \u002F>\u003Cpath d=\"M440 112 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Cpath d=\"M560 112 H571\" class=\"d-line\" \u002F>\u003Cpath d=\"M580 112 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Cpath d=\"M80 154 V168\" class=\"d-dash\" \u002F>\u003Ctext x=\"80\" y=\"186\" text-anchor=\"middle\" class=\"d-label\">human gate\u003C\u002Ftext>\u003Cpath d=\"M220 154 V168\" class=\"d-dash\" \u002F>\u003Ctext x=\"220\" y=\"186\" text-anchor=\"middle\" class=\"d-label\">human gate\u003C\u002Ftext>\u003Cpath d=\"M640 154 V168\" class=\"d-dash\" \u002F>\u003Ctext x=\"640\" y=\"186\" text-anchor=\"middle\" class=\"d-label\">human gate\u003C\u002Ftext>\u003Ctext x=\"360\" y=\"240\" text-anchor=\"middle\" class=\"d-small\">Stop at a gate: a short spec or plan is cheap to fix, a merged change is not.\u003C\u002Ftext>",[201],"The gates sit under the spec, the plan and the verification. A short spec or plan is cheap to fix, a merged change is not.",{"type":155,"ordered":156,"items":203},[204,209,214,219,224],[205,208],{"tag":160,"children":206},[207],"Spec."," The problem, the behaviour you want, the non-goals and the acceptance criteria. A person approves it.",[210,213],{"tag":160,"children":211},[212],"Plan."," The files and interfaces that change, the order of work, the risks and what is out of scope. A person approves it, because a wrong design is cheapest to catch here.",[215,218],{"tag":160,"children":216},[217],"Tasks."," A checklist of small steps, each with its own check.",[220,223],{"tag":160,"children":221},[222],"Implementation."," The agent writes code and tests for one task at a time and ticks it off only when the check passes.",[225,228],{"tag":160,"children":226},[227],"Verification."," Each criterion maps to a test, a command or a manual check, and the evidence is attached to the change. A person reviews that evidence, not the diff on its own.",{"type":142,"content":230},[231],"Anthropic's best-practice guide for Claude Code describes the same order: explore, plan, implement, commit. It says planning is most useful when you are uncertain about the approach, when the change modifies several files, or when you are unfamiliar with the code, and that if you could describe the diff in one sentence, you should skip the plan. I agree with both halves.",{"type":149,"level":150,"id":117,"text":118},{"type":142,"content":234},[235],"A spec is short. If it runs past a few hundred words, it is probably describing the implementation, which belongs in the plan. The acceptance criteria matter most, and I would start from EARS, the Easy Approach to Requirements Syntax. Alistair Mavin and colleagues at Rolls-Royce devised it while analysing airworthiness rules for a jet engine control system, and it was first published in 2009. Each requirement takes one of a few shapes: WHEN for events, WHILE for states, IF and THEN for unwanted situations, and WHERE for optional features.",{"type":237,"code":238},"code","# Spec: cart login prompt and payment safety (example)\n\n## Problem\nLogged-out visitors reach checkout without learning that an account unlocks the B2B price list. A timeout on the payment call sometimes leads to a second order.\n\n## Behaviour\nWhat the shop must do, from the user's point of view.\n\n## Non-goals\n- No change to how price lists are modelled\n- No redesign of the cart layout\n\n## Acceptance criteria\n- AC-1: WHEN a logged-out visitor opens the cart, the system SHALL show a login prompt above the checkout button.\n- AC-2: WHILE the price cache is older than 15 minutes, the system SHALL refetch prices before it shows totals.\n- AC-3: IF the payment call times out, THEN the system SHALL keep the order pending and SHALL NOT submit a second charge.\n- AC-4: WHERE the account has a negotiated price list, the system SHALL show that list instead of the public price.\n\n## Verification\n- AC-1 -> tests\u002Fcart\u002Flogin-prompt.spec.ts (e2e)\n- AC-2 -> tests\u002Fpricing\u002Fcache-refetch.test.ts (unit, fake clock)\n- AC-3 -> tests\u002Fcheckout\u002Ftimeout.spec.ts (e2e) and a payment unit test\n- AC-4 -> tests\u002Fpricing\u002Fnegotiated.test.ts (unit)\nDone when every line above passes in CI and the evidence is attached to the change.",{"type":142,"content":240},[241],"The criteria in this example are invented for a B2B shop. The shape is what matters: one sentence with a trigger or a state, so a test can be written from it, and an ID, so the plan and the evidence can point at it. The IF and THEN line matters most, because it protects against a double charge, and it is the kind of case an agent does not think of unless you write it down.",{"type":149,"level":150,"id":120,"text":121},{"type":142,"content":244},[245],"The plan answers how, and it should be short enough to review in minutes. Anthropic's guidance describes the most useful specs as ones that name the files and interfaces involved, state what is out of scope, and end with an end-to-end check that proves the feature works. The plan should also list the known risks and the order of work. Build the riskiest part first, while the design can still change.",{"type":237,"code":247},"# Plan: cart login prompt and payment safety\n\n## Files that change\n- src\u002Fcart\u002FCartPage.vue: login prompt for logged-out visitors (AC-1)\n- src\u002Fpricing\u002FpriceCache.ts: refetch after 15 minutes (AC-2)\n- src\u002Fpricing\u002Fnegotiated.ts: negotiated list lookup (AC-4)\n- src\u002Fcheckout\u002Fpayment.ts: idempotency key, timeout keeps the order pending (AC-3)\n\n## Interfaces\n- payment.charge() gains an idempotencyKey argument; both callers are updated\n- priceCache.get() keeps its signature and also returns fetchedAt\n\n## Order of work\n1. Idempotency key and timeout path (riskiest, built first)\n2. Negotiated price lookup\n3. Price cache refetch\n4. Login prompt (UI, last)\n\n## Risks\n- A retry after a timeout could charge twice (covered by AC-3)\n\n## Out of scope\n- Changing how price lists are stored",{"type":142,"content":249},[250,251,256],"The task list makes the agent's work visible. Each task is small enough that its check can fail for one reason, and it names the criterion it serves. A tick without evidence is only a claim, and claims are what vibe coding runs on. The same principle sits behind the guides and sensors in my ",{"tag":252,"to":253,"children":254},"link","\u002Fblog\u002Fharness-engineering-coding-agents",[255],"harness engineering article",".",{"type":237,"code":258},"# Tasks: cart login prompt and payment safety\n\n- [x] T1 Idempotency key on payment.charge() (AC-3)\n      check: unit test 'retry after timeout does not charge twice' passes\n- [x] T2 Timeout path keeps the order pending (AC-3)\n      check: e2e 'payment timeout' passes\n- [ ] T3 Negotiated price lookup (AC-4)\n      check: unit test 'account sees negotiated list' passes\n- [ ] T4 Price cache refetch after 15 minutes (AC-2)\n      check: unit test with fake clock passes\n- [ ] T5 Login prompt for logged-out visitors (AC-1)\n      check: e2e 'logged-out cart' passes\n- [ ] T6 Full suite green; attach output and map AC-1 to AC-4 to tests",{"type":149,"level":150,"id":123,"text":124},{"type":142,"content":261},[262],"The tools differ less in their workflow than in where the spec lives and what enforces the gate. The table lists what I could confirm in the official documentation in October 2026.",{"type":264,"head":265,"rows":274},"table",[266,268,270,272],[267],"Tool",[269],"What it keeps",[271],"Where the human gate is",[273],"Caveat",[275,283,292,301,310],[276,277,279,281],[19],[278],"Chat commands for constitution, specify, plan, tasks, implement and converge, plus a spec folder per feature",[280],"Chat steps; the constitution requires tests to be approved before implementation",[282],"Heavy paperwork, as Böckeler found",[284,286,288,290],[285],"Kiro specs",[287],"requirements.md (or bugfix.md), design.md and tasks.md for each spec",[289],"Quick Spec generates all three files in one pass without approval gates",[291],"Ceremony a small bug does not need",[293,295,297,299],[294],"Claude Code plan mode",[296],"The plan in the session; Ctrl+G opens it in your editor",[298],"You approve the plan, or press Shift+Tab, to leave plan mode",[300],"Read-only planning, not a durable spec",[302,304,306,308],[303],"Cline Plan and Act",[305],"The plan stays in the chat unless you ask for a markdown summary",[307],"Plan mode cannot edit files or run commands; Act mode can",[309],"No approval setting for the switch is described",[311,313,315,317],[312],"Codex \u002Fplan",[314],"Listed as “Toggle plan mode for multi-step planning” in OpenAI's command reference",[316],"Not described in the pages I could open",[318],"Read-only behaviour and storage unverified, so I do not rely on it",{"type":142,"content":320},[321],"Two rows in that table matter more than the rest. A plan mode is a mode, not a spec: it stops the agent from editing while it reasons, which is valuable, but it leaves you without a durable artefact to review, diff or test against. The file-based tools move the cost into review, which I come back to below.",{"type":142,"content":323},[324],"Kiro's introduction says its user stories carry EARS acceptance criteria, the shape I recommend below. What matters is a fixed sentence shape that a test can be written from. Spec Kit runs its steps in the agent's chat, with the setup in the terminal:",{"type":237,"code":326},"uv tool install specify-cli\nspecify init my-project --integration copilot\ncd my-project",{"type":142,"content":328},[329],"Then, in the chat, run \u002Fspeckit-constitution once per project and \u002Fspeckit-specify, \u002Fspeckit-plan, \u002Fspeckit-tasks and \u002Fspeckit-implement for each feature. The README lists Python 3.11 or newer, uv and a supported AI coding agent as prerequisites, and its examples use GitHub Copilot. In the methodology document, the constitution's test-first article requires that tests are approved before implementation.",{"type":149,"level":150,"id":126,"text":127},{"type":142,"content":332},[333],"A spec changes what the reviewer does. Without one, the reviewer reconstructs the intent from the diff. With one, the reviewer asks four concrete questions: is every criterion covered by a check that can fail, is there evidence that each check passes, did anything outside the scope change, and does the code still match the risks the plan named? Each criterion points at a task, each task at a test, and each test leaves evidence behind.",{"type":194,"attrs":335,"inner":338,"caption":339},{"viewBox":336,"role":197,"aria-labelledby":337},"0 0 720 210","d2-trace-t d2-trace-d","\u003Ctitle id=\"d2-trace-t\">From acceptance criterion to evidence\u003C\u002Ftitle>\u003Cdesc id=\"d2-trace-d\">Four boxes in a row: an acceptance criterion, the task that implements it, the test that checks it and the continuous integration output that shows the result. A reviewer compares the evidence with the criterion, not with the diff alone.\u003C\u002Fdesc>\u003Ctext x=\"20\" y=\"24\" class=\"d-label\">Traceability for one criterion\u003C\u002Ftext>\u003Crect x=\"12\" y=\"60\" width=\"144\" height=\"80\" rx=\"10\" class=\"d-gold\" \u002F>\u003Ctext x=\"84\" y=\"94\" text-anchor=\"middle\" class=\"d-text\">AC-3\u003C\u002Ftext>\u003Ctext x=\"84\" y=\"116\" text-anchor=\"middle\" class=\"d-small\">no double charge\u003C\u002Ftext>\u003Crect x=\"196\" y=\"60\" width=\"144\" height=\"80\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"268\" y=\"94\" text-anchor=\"middle\" class=\"d-text\">Task T2\u003C\u002Ftext>\u003Ctext x=\"268\" y=\"116\" text-anchor=\"middle\" class=\"d-small\">timeout path\u003C\u002Ftext>\u003Crect x=\"380\" y=\"60\" width=\"144\" height=\"80\" rx=\"10\" class=\"d-sky\" \u002F>\u003Ctext x=\"452\" y=\"94\" text-anchor=\"middle\" class=\"d-text\">e2e test\u003C\u002Ftext>\u003Ctext x=\"452\" y=\"116\" text-anchor=\"middle\" class=\"d-small\">payment timeout\u003C\u002Ftext>\u003Crect x=\"564\" y=\"60\" width=\"144\" height=\"80\" rx=\"10\" class=\"d-mint\" \u002F>\u003Ctext x=\"636\" y=\"94\" text-anchor=\"middle\" class=\"d-text\">CI output\u003C\u002Ftext>\u003Ctext x=\"636\" y=\"116\" text-anchor=\"middle\" class=\"d-small\">exit code 0\u003C\u002Ftext>\u003Cpath d=\"M156 100 H187\" class=\"d-line\" \u002F>\u003Cpath d=\"M196 100 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Cpath d=\"M340 100 H371\" class=\"d-line\" \u002F>\u003Cpath d=\"M380 100 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Cpath d=\"M524 100 H555\" class=\"d-line\" \u002F>\u003Cpath d=\"M564 100 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Cpath d=\"M84 146 V168 H636 V146\" class=\"d-dash\" \u002F>\u003Ctext x=\"360\" y=\"192\" text-anchor=\"middle\" class=\"d-small\">The reviewer compares the evidence with the criterion.\u003C\u002Ftext>",[340],"A criterion without a test, or a test without evidence, is a gap the reviewer should flag.",{"type":142,"content":342},[343],"Anthropic's guidance makes the same point from the agent's side. Give it a check it can run, such as tests, a build or a screenshot to compare, and ask for evidence rather than assertions: the test output, the command it ran and what it returned. The same guide suggests a second opinion, a fresh subagent that reviews the diff against the plan.",{"type":183,"variant":345,"title":346,"body":347},"tip","Let a fresh reviewer check the diff against the plan",[348],[349],"Ask a subagent, in a fresh context, to check that every planned requirement is implemented, that the listed edge cases have tests and that nothing outside the task's scope changed. Tell it to flag only gaps that affect correctness or the stated requirements, because a reviewer prompted to find gaps reports some even when the work is sound.",{"type":142,"content":351},[352,353,357],"Agent-written pull requests make this more urgent. My article on the ",{"tag":252,"to":354,"children":355},"\u002Fblog\u002Fai-generated-pr-review-bottleneck",[356],"AI code review bottleneck"," covers what happens when the review queue fills up, and evidence-based review is one way to keep it moving.",{"type":142,"content":359},[360],"Keep examples in specs synthetic. Specs and plans get committed, reviewed and often sent to a model as context, so a real customer name in an example becomes personal data processed by a third party. If GDPR applies, find out where your agent provider processes that context and how long it is kept, and check that your data processing agreement covers it.",{"type":149,"level":150,"id":129,"text":130},{"type":142,"content":363},[364],"The first cost is human time, not tokens. Writing and reviewing a spec and a plan takes real time, and that is the price of the approach. I have no measured figure for what it saves, and vendor figures deserve the same scrutiny as the METR numbers. The token side is small and easy to check.",{"type":142,"content":366},[367],"At Anthropic's current prices, a spec of about 6,000 tokens that the agent reads on each of 40 calls costs about 48 cents uncached on Claude Sonnet 5.5, at $2 per million input tokens. With the 5-minute prompt cache, the first write costs $2.50 per million and each later read costs $0.10 per million, so the same 40 calls come to about 4 cents if they arrive within the cache window. These are my own calculations from the official pricing page. They leave out the code the agent reads and every output token, so they show the spec's share of the bill, not the bill.",{"type":142,"content":369},[370],"Two caveats apply. The newer tokenizer used by Claude 4.7 and later produces about 30% more tokens for the same text, so count the spec in tokens, not words. Attention is the bigger cost, though. Anthropic's guidance says performance degrades as the context window fills, and that the model may start to forget earlier instructions. Keep the spec to the criteria, the constraints and the non-goals, and let the tests carry the detail.",{"type":264,"head":372,"rows":379},[373,375,377],[374],"Change",[376],"Approach",[378],"Why",[380,387,394,401,408],[381,383,385],[382],"A typo, a log line or a rename",[384],"Prompt directly, no plan",[386],"The whole diff fits in one sentence",[388,390,392],[389],"A bug with a known cause",[391],"Failing test first, then the fix",[393],"The failing test is the smallest useful spec",[395,397,399],[396],"A feature inside one module",[398],"Short plan in plan mode, checked at the end",[400],"Review is cheap, and planning still catches a wrong approach",[402,404,406],[403],"A change across modules, or a risky path such as payments or auth",[405],"Full spec, plan, tasks and evidence",[407],"Rework costs more than the paperwork",[409,411,413],[410],"Code that others will extend for months",[412],"A spec kept alive as documentation",[414],"A stale spec misleads more than no spec",{"type":142,"content":416},[417],"Böckeler found that spec-kit “felt like overkill for the size of the problem” on a mid-sized feature, and she argued that a useful tool has to support several workflow sizes. The process should follow the size of the change, not the tool you already have.",{"type":149,"level":150,"id":132,"text":133},{"type":155,"ordered":156,"items":420},[421,426,431],[422,425],{"tag":160,"children":423},[424],"Paperwork without review."," Böckeler found that spec-kit “created a LOT of markdown files for me to review”, and she would rather review code than those files. If nobody reads the spec, it is theatre.",[427,430],{"tag":160,"children":428},[429],"Specs that go stale."," Böckeler separates spec-first, spec-anchored (the spec survives and maintains the feature) and spec-as-source. Two of the three tools she reviewed are spec-first, and many approaches stay vague about how the spec is kept up to date.",[432,435],{"tag":160,"children":433},[434],"Over-elaboration."," Thoughtworks placed spec-driven development in its Assess ring in November 2025, noting that its workflows “remain elaborate and opinionated”, and warned that “we may be relearning a bitter lesson”.",{"type":142,"content":437},[438],"A spec does not make the agent correct. It makes mistakes visible sooner and gives the reviewer something to check. Böckeler also flagged non-determinism: in Tessl, generating code repeatedly from the same spec does not give the same result. The honest summary is that spec-driven development is discipline for the parts of a change that can go wrong, and overhead for the rest.",{"type":149,"level":150,"id":135,"text":136},{"type":142,"content":441},[442,443,447,448,452],"None of this needs a particular tool. A markdown file, a checklist and a test suite are enough to start. My article on ",{"tag":252,"to":444,"children":445},"\u002Fblog\u002Fcoding-agent-skills-workflow",[446],"coding agent skills"," shows the same discipline applied to a bug, from report to pull request, and my note on ",{"tag":252,"to":449,"children":450},"\u002Fblog\u002Fhuman-in-the-loop-ai-agents",[451],"human in the loop agents"," covers where the approval gates belong.",{"type":155,"ordered":454,"items":455},true,[456,458,460,462,464,466],[457],"Pick one change next week that touches two or more modules, and write its spec with acceptance criteria before you open the agent.",[459],"Write each criterion as one sentence in a fixed shape, such as EARS, and give it an ID that the plan and the tasks can reference.",[461],"Ask for a plan, read it for ten minutes, and change the file list and the order of work before any code is written.",[463],"Make every task tick depend on evidence in the change: a test name, a command with its exit code, or a screenshot.",[465],"Delete the parts of the spec that the code no longer matches, rather than letting them drift.",[467],"Judge the result by rework and review time, not by how fast the first draft appeared. The METR results show how unreliable that feeling is.",{"type":149,"level":150,"id":138,"text":139},{"type":155,"ordered":454,"items":470},[471,475,478,481,484,487,490,493,496,499,502,505,508,511,514],[472],{"tag":473,"href":38,"children":474},"a",[37],[476],{"tag":473,"href":28,"children":477},[40],[479],{"tag":473,"href":43,"children":480},[42],[482],{"tag":473,"href":46,"children":483},[45],[485],{"tag":473,"href":49,"children":486},[48],[488],{"tag":473,"href":31,"children":489},[51],[491],{"tag":473,"href":54,"children":492},[53],[494],{"tag":473,"href":57,"children":495},[56],[497],{"tag":473,"href":60,"children":498},[59],[500],{"tag":473,"href":63,"children":501},[62],[503],{"tag":473,"href":66,"children":504},[65],[506],{"tag":473,"href":69,"children":507},[68],[509],{"tag":473,"href":26,"children":510},[71],[512],{"tag":473,"href":74,"children":513},[73],[515],{"tag":473,"href":77,"children":516},[76],[518,599,654,733],{"slug":519,"published":520,"minutes":6,"category":7,"tags":521,"keywords":527,"about":536,"sources":546,"cover":593,"og":594,"expertise":80,"locales":595,"lang":82,"title":596,"description":597,"coverAlt":598},"ai-assisted-development-economics","2026-10-01",[522,523,524,525,526],"AI coding agents","Developer productivity","Engineering economics","Team cost","GDPR",[528,529,530,531,532,533,534,535],"AI coding agents cost","developer productivity AI study","METR AI developer slowdown","AI coding break-even cost model","one senior engineer vs team","AI coding agent vs agency","DORA AI adoption report","GDPR processor contract AI coding tools",[537,540,543],{"name":538,"url":539},"General Data Protection Regulation","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FGeneral_Data_Protection_Regulation",{"name":541,"url":542},"Randomized controlled trial","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FRandomized_controlled_trial",{"name":544,"url":545},"Bus factor","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FBus_factor",[547,549,552,554,557,560,563,566,569,572,575,578,581,584,587,590],{"title":548,"url":63},"METR: early-2025 AI and experienced open-source developers",{"title":550,"url":551},"Becker et al.: arXiv 2507.09089","https:\u002F\u002Farxiv.org\u002Fabs\u002F2507.09089",{"title":553,"url":66},"METR: developer productivity experiment design, February 2026",{"title":555,"url":556},"Peng et al.: GitHub Copilot controlled experiment, arXiv 2302.06590","https:\u002F\u002Farxiv.org\u002Fhtml\u002F2302.06590v1",{"title":558,"url":559},"Cui et al.: three field experiments with software developers","https:\u002F\u002Fwww.microsoft.com\u002Fen-us\u002Fresearch\u002F?p=1148213",{"title":561,"url":562},"DORA: State of AI-assisted Software Development 2025","https:\u002F\u002Fresearch.google\u002Fpubs\u002Fdora-2025-state-of-ai-assisted-software-development-report\u002F",{"title":564,"url":565},"Google Cloud: highlights from the 2024 DORA report","https:\u002F\u002Fcloud.google.com\u002Fblog\u002Fproducts\u002Fdevops-sre\u002Fannouncing-the-2024-dora-report",{"title":567,"url":568},"Stack Overflow: Developer Survey 2025, AI","https:\u002F\u002Fsurvey.stackoverflow.co\u002F2025\u002Fai",{"title":570,"url":571},"Ziftci et al.: Migrating Code At Scale With LLMs At Google","https:\u002F\u002Farxiv.org\u002Fabs\u002F2504.09691",{"title":573,"url":574},"Alshahwan et al.: Automated Unit Test Improvement using Large Language Models at Meta","https:\u002F\u002Farxiv.org\u002Fabs\u002F2402.09171",{"title":576,"url":577},"GitClear: AI Copilot Code Quality: 2025 Look Back at 12 Months of Data","https:\u002F\u002Fwww.gitclear.com\u002Fai_assistant_code_quality_2025_research",{"title":579,"url":580},"GDPR, Regulation (EU) 2016\u002F679, Article 28","https:\u002F\u002Feur-lex.europa.eu\u002Feli\u002Freg\u002F2016\u002F679\u002Foj\u002Feng",{"title":582,"url":583},"Anthropic: Commercial Terms of Service","https:\u002F\u002Fwww.anthropic.com\u002Flegal\u002Fcommercial-terms",{"title":585,"url":586},"Claude: pricing","https:\u002F\u002Fclaude.com\u002Fpricing",{"title":588,"url":589},"Claude Platform: data residency","https:\u002F\u002Fplatform.claude.com\u002Fdocs\u002Fen\u002Fbuild-with-claude\u002Fdata-residency",{"title":591,"url":592},"GitHub Copilot: plans and pricing","https:\u002F\u002Fgithub.com\u002Ffeatures\u002Fcopilot\u002Fplans","\u002Fimages\u002Fblog\u002Fai-assisted-development-economics\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fai-assisted-development-economics\u002Fog.jpg",[82,83,84],"One senior with coding agents versus a team: what the evidence says","The METR, DORA, Microsoft and GitHub studies on AI coding tools, what they do not prove, and a break-even cost model for one senior versus a team or agency.","Cover art for AI coding economics: a draft-to-ship pipeline where agents speed up drafting, and review and quality costs take part of the gain back.",{"slug":600,"published":601,"minutes":602,"category":7,"tags":603,"keywords":609,"about":619,"sources":629,"cover":648,"og":649,"expertise":80,"locales":650,"lang":82,"title":651,"description":652,"coverAlt":653},"mcp-tool-design-lessons-jira-server","2026-09-18",10,[604,605,606,607,608],"MCP","Tool design","Context engineering","Jira","Agents",[610,611,612,613,614,615,616,617,618],"MCP tool design","MCP best practices","MCP tool descriptions","agent tool selection","MCP context bloat","how many tools should an MCP server have","MCP tool definition token cost","MCP error handling isError","Jira MCP server",[620,623,626],{"name":621,"url":622},"Model Context Protocol","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FModel_Context_Protocol",{"name":624,"url":625},"Jira (software)","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FJira_(software)",{"name":627,"url":628},"Intelligent agent","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FIntelligent_agent",[630,633,636,639,642,645],{"title":631,"url":632},"Writing effective tools for agents – with agents (Anthropic)","https:\u002F\u002Fwww.anthropic.com\u002Fengineering\u002Fwriting-tools-for-agents",{"title":634,"url":635},"Introducing advanced tool use on the Claude Developer Platform (Anthropic)","https:\u002F\u002Fwww.anthropic.com\u002Fengineering\u002Fadvanced-tool-use",{"title":637,"url":638},"Code execution with MCP (Anthropic)","https:\u002F\u002Fwww.anthropic.com\u002Fengineering\u002Fcode-execution-with-mcp",{"title":640,"url":641},"MCP vs CLI: context window cost (Blocks.ai)","https:\u002F\u002Fblocks.ai\u002Fblog\u002Fmcp-vs-cli-context-window-cost",{"title":643,"url":644},"Demystifying evals for AI agents (Anthropic)","https:\u002F\u002Fwww.anthropic.com\u002Fengineering\u002Fdemystifying-evals-for-ai-agents",{"title":646,"url":647},"MCP 2026-07-28 specification: Tools","https:\u002F\u002Fmodelcontextprotocol.io\u002Fspecification\u002F2026-07-28\u002Fserver\u002Ftools","\u002Fimages\u002Fblog\u002Fmcp-tool-design-lessons-jira-server\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fmcp-tool-design-lessons-jira-server\u002Fog.jpg",[82,83,84],"MCP tool design: lessons from a 20-tool Jira server","MCP tool design that agents get right: token cost of tool definitions, when to merge tools, naming, concise output, errors that steer and a small selection eval.","Network diagram with a Jira MCP server at the hub and five satellites: search, create, transition, comments and test runs",{"slug":655,"published":656,"minutes":657,"category":7,"tags":658,"keywords":661,"about":672,"sources":678,"cover":727,"og":728,"expertise":80,"locales":729,"lang":82,"title":730,"description":731,"coverAlt":732},"ai-agent-memory-design","2026-09-10",13,[659,606,660,526],"AI agent memory","Memory poisoning",[662,663,664,665,666,667,668,669,670,671],"AI agent memory design","long-term memory for AI agents","episodic semantic procedural memory LLM","agent memory architecture","ChatGPT memory vs Claude memory","Claude memory tool","AI memory poisoning","LLM context compaction","AI agent memory GDPR","short-term vs long-term memory agents",[673,674,675],{"name":627,"url":628},{"name":538,"url":539},{"name":676,"url":677},"Prompt injection","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FPrompt_injection",[679,682,685,688,691,694,697,700,703,706,709,712,715,718,721,724],{"title":680,"url":681},"Anthropic docs: Memory tool","https:\u002F\u002Fplatform.claude.com\u002Fdocs\u002Fen\u002Fagents-and-tools\u002Ftool-use\u002Fmemory-tool",{"title":683,"url":684},"Anthropic docs: Context editing","https:\u002F\u002Fplatform.claude.com\u002Fdocs\u002Fen\u002Fbuild-with-claude\u002Fcontext-editing",{"title":686,"url":687},"Anthropic Engineering: Effective context engineering for AI agents","https:\u002F\u002Fwww.anthropic.com\u002Fengineering\u002Feffective-context-engineering-for-ai-agents",{"title":689,"url":690},"Sumers et al.: Cognitive Architectures for Language Agents (CoALA)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2309.02427",{"title":692,"url":693},"Packer et al.: MemGPT, Towards LLMs as Operating Systems","https:\u002F\u002Farxiv.org\u002Fabs\u002F2310.08560",{"title":695,"url":696},"Park et al.: Generative Agents, Interactive Simulacra of Human Behavior","https:\u002F\u002Farxiv.org\u002Fabs\u002F2304.03442",{"title":698,"url":699},"Unit 42: When AI Remembers Too Much, persistent behaviors in agents memory","https:\u002F\u002Funit42.paloaltonetworks.com\u002Findirect-prompt-injection-poisons-ai-longterm-memory\u002F",{"title":701,"url":702},"From Untrusted Input to Trusted Memory: A Systematic Study of Memory Poisoning Attacks in LLM Agents (preprint)","https:\u002F\u002Farxiv.org\u002Fhtml\u002F2606.04329v1",{"title":704,"url":705},"The Hacker News: ChatGPT macOS flaw could have enabled long-term spyware via memory function","https:\u002F\u002Fthehackernews.com\u002F2024\u002F09\u002Fchatgpt-macos-flaw-couldve-enabled-long.html",{"title":707,"url":708},"Vectorize: OWASP ASI06, Memory and Context Poisoning explained","https:\u002F\u002Fvectorize.io\u002Farticles\u002Fowasp-asi06",{"title":710,"url":711},"Claude Help Center: Use chat search and memory to build on previous context","https:\u002F\u002Fsupport.claude.com\u002Fen\u002Farticles\u002F11817273-use-claude-s-chat-search-and-memory-to-build-on-previous-context",{"title":713,"url":714},"OpenAI Help Center: Memory in ChatGPT","https:\u002F\u002Fhelp.openai.com\u002Fen\u002Farticles\u002F8590148-memory-faq",{"title":716,"url":717},"OpenAI Help Center: Dots privacy, security, and safety FAQs","https:\u002F\u002Fhelp.openai.com\u002Fen\u002Farticles\u002F20001529-dots-privacy-security-and-safety-faqs",{"title":719,"url":720},"Flavio Copes: A deep dive into OpenAI dots (quotes the dots documentation on memory)","https:\u002F\u002Fflaviocopes.com\u002Fopenai-dots\u002F",{"title":722,"url":723},"GDPR Article 5: Principles relating to processing of personal data","https:\u002F\u002Fgdpr-info.eu\u002Fart-5-gdpr\u002F",{"title":725,"url":726},"GDPR Article 17: Right to erasure","https:\u002F\u002Fgdpr-info.eu\u002Fart-17-gdpr\u002F","\u002Fimages\u002Fblog\u002Fai-agent-memory-design\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fai-agent-memory-design\u002Fog.jpg",[82,83,84],"Designing memory for AI agents: tiers, write rules, poisoning and GDPR","How to design AI agent memory: context vs session vs long-term tiers, what to write and never store, retrieval, compaction, poisoning and GDPR erasure.","Diagram: nested memory layers of an AI agent, from the working context window through session state to long-term episodic and semantic memory.",{"slug":734,"published":735,"minutes":736,"category":7,"tags":737,"keywords":742,"about":752,"sources":761,"cover":786,"og":787,"expertise":80,"locales":788,"lang":82,"title":789,"description":790,"coverAlt":791},"harness-engineering-coding-agents","2026-09-04",8,[738,10,739,740,741],"Harness engineering","Code quality","Mutation testing","TDD",[743,744,745,746,747,748,749,750,751],"harness engineering","harness engineering coding agents","AI code quality","coding agent guardrails","mutation testing AI generated tests","red\u002Fgreen TDD with AI agents","guides and sensors coding agents","how to make AI agent pull requests mergeable","can I trust tests written by AI",[753,755,758],{"name":740,"url":754},"https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FMutation_testing",{"name":756,"url":757},"Test-driven development","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FTest-driven_development",{"name":759,"url":760},"Static program analysis","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FStatic_program_analysis",[762,765,768,771,774,777,780,783],{"title":763,"url":764},"Birgitta Böckeler: Harness engineering for coding agent users (Apr 2026)","https:\u002F\u002Fmartinfowler.com\u002Farticles\u002Fharness-engineering.html",{"title":766,"url":767},"Birgitta Böckeler: Maintainability sensors for coding agents (May 2026)","https:\u002F\u002Fmartinfowler.com\u002Farticles\u002Fsensors-for-coding-agents.html",{"title":769,"url":770},"Simon Willison: Agentic Engineering Patterns","https:\u002F\u002Fsimonwillison.net\u002Fguides\u002Fagentic-engineering-patterns\u002F",{"title":772,"url":773},"Simon Willison: First run the tests","https:\u002F\u002Fsimonwillison.net\u002Fguides\u002Fagentic-engineering-patterns\u002Ffirst-run-the-tests\u002F",{"title":775,"url":776},"Anthropic: Effective harnesses for long-running agents (Nov 2025)","https:\u002F\u002Fwww.anthropic.com\u002Fengineering\u002Feffective-harnesses-for-long-running-agents",{"title":778,"url":779},"Claude Code docs: How Claude remembers your project","https:\u002F\u002Fcode.claude.com\u002Fdocs\u002Fen\u002Fmemory",{"title":781,"url":782},"Stryker Mutator documentation","https:\u002F\u002Fstryker-mutator.io\u002Fdocs\u002F",{"title":784,"url":785},"Infection: command line options","https:\u002F\u002Finfection.github.io\u002Fguide\u002Fcommand-line-options.html","\u002Fimages\u002Fblog\u002Fharness-engineering-coding-agents\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fharness-engineering-coding-agents\u002Fog.jpg",[82,83,84],"Harness engineering: guides and sensors that make agent PRs mergeable","Harness engineering for coding agents: guides and sensors, where to run each check, red\u002Fgreen TDD, and mutation testing to verify the tests the agent wrote.","Concentric rings around a coding agent's model: behaviour, architecture fitness and maintainability harnesses, from outside in.",1791636875225]