[{"data":1,"prerenderedAt":809},["ShallowReactive",2],{"blog-ai-assisted-development-economics-en":3},{"slug":4,"published":5,"minutes":6,"category":7,"tags":8,"keywords":14,"about":23,"sources":33,"cover":82,"og":83,"expertise":84,"locales":85,"lang":86,"title":89,"description":90,"coverAlt":91,"metaTitle":92,"takeaways":93,"faq":99,"toc":112,"blocks":140,"others":538},"ai-assisted-development-economics","2026-10-01",11,"agents",[9,10,11,12,13],"AI coding agents","Developer productivity","Engineering economics","Team cost","GDPR",[15,16,17,18,19,20,21,22],"AI coding agents cost","developer productivity AI study","METR AI developer slowdown","AI coding break-even cost model","one senior engineer vs team","AI coding agent vs agency","DORA AI adoption report","GDPR processor contract AI coding tools",[24,27,30],{"name":25,"url":26},"General Data Protection Regulation","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FGeneral_Data_Protection_Regulation",{"name":28,"url":29},"Randomized controlled trial","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FRandomized_controlled_trial",{"name":31,"url":32},"Bus factor","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FBus_factor",[34,37,40,43,46,49,52,55,58,61,64,67,70,73,76,79],{"title":35,"url":36},"METR: early-2025 AI and experienced open-source developers","https:\u002F\u002Fmetr.org\u002Fblog\u002F2025-07-10-early-2025-ai-experienced-os-dev-study\u002F",{"title":38,"url":39},"Becker et al.: arXiv 2507.09089","https:\u002F\u002Farxiv.org\u002Fabs\u002F2507.09089",{"title":41,"url":42},"METR: developer productivity experiment design, February 2026","https:\u002F\u002Fmetr.org\u002Fblog\u002F2026-02-24-uplift-update\u002F",{"title":44,"url":45},"Peng et al.: GitHub Copilot controlled experiment, arXiv 2302.06590","https:\u002F\u002Farxiv.org\u002Fhtml\u002F2302.06590v1",{"title":47,"url":48},"Cui et al.: three field experiments with software developers","https:\u002F\u002Fwww.microsoft.com\u002Fen-us\u002Fresearch\u002F?p=1148213",{"title":50,"url":51},"DORA: State of AI-assisted Software Development 2025","https:\u002F\u002Fresearch.google\u002Fpubs\u002Fdora-2025-state-of-ai-assisted-software-development-report\u002F",{"title":53,"url":54},"Google Cloud: highlights from the 2024 DORA report","https:\u002F\u002Fcloud.google.com\u002Fblog\u002Fproducts\u002Fdevops-sre\u002Fannouncing-the-2024-dora-report",{"title":56,"url":57},"Stack Overflow: Developer Survey 2025, AI","https:\u002F\u002Fsurvey.stackoverflow.co\u002F2025\u002Fai",{"title":59,"url":60},"Ziftci et al.: Migrating Code At Scale With LLMs At Google","https:\u002F\u002Farxiv.org\u002Fabs\u002F2504.09691",{"title":62,"url":63},"Alshahwan et al.: Automated Unit Test Improvement using Large Language Models at Meta","https:\u002F\u002Farxiv.org\u002Fabs\u002F2402.09171",{"title":65,"url":66},"GitClear: AI Copilot Code Quality: 2025 Look Back at 12 Months of Data","https:\u002F\u002Fwww.gitclear.com\u002Fai_assistant_code_quality_2025_research",{"title":68,"url":69},"GDPR, Regulation (EU) 2016\u002F679, Article 28","https:\u002F\u002Feur-lex.europa.eu\u002Feli\u002Freg\u002F2016\u002F679\u002Foj\u002Feng",{"title":71,"url":72},"Anthropic: Commercial Terms of Service","https:\u002F\u002Fwww.anthropic.com\u002Flegal\u002Fcommercial-terms",{"title":74,"url":75},"Claude: pricing","https:\u002F\u002Fclaude.com\u002Fpricing",{"title":77,"url":78},"Claude Platform: data residency","https:\u002F\u002Fplatform.claude.com\u002Fdocs\u002Fen\u002Fbuild-with-claude\u002Fdata-residency",{"title":80,"url":81},"GitHub Copilot: plans and pricing","https:\u002F\u002Fgithub.com\u002Ffeatures\u002Fcopilot\u002Fplans","\u002Fimages\u002Fblog\u002Fai-assisted-development-economics\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fai-assisted-development-economics\u002Fog.jpg","ai-engineer",[86,87,88],"en","de","hu","One senior with coding agents versus a team: what the evidence says","The METR, DORA, Microsoft and GitHub studies on AI coding tools, what they do not prove, and a break-even cost model for one senior versus a team or agency.","Cover art for AI coding economics: a draft-to-ship pipeline where agents speed up drafting, and review and quality costs take part of the gain back.","Coding agents vs a team: the evidence · Balázs Csorba",[94,95,96,97,98],"The studies disagree, and the pattern is useful: in a 2025 trial, experienced developers on codebases they knew took 19% longer with AI, while in a 2023 trial a single bounded task took 55.8% less time with Copilot.","Measure the net gain end to end, from ticket to production at the same quality. Speed gained at the draft can be spent again in review, rework and incidents.","The break-even is simple: the net gain has to beat the tool, review and quality costs, expressed as a share of the engineer’s loaded cost.","One senior with agents concentrates risk in one person and one vendor. Price in a second reviewer, a written runbook and a fallback before you decide.","Sign a data processing agreement and confirm training and processing region in writing before client data goes into a prompt, as Article 28 of the GDPR requires for processors.",[100,103,106,109],{"q":101,"a":102},"Do AI coding agents make experienced developers faster?","The evidence does not settle it. A 2025 randomised trial by METR found that experienced open-source developers took 19% longer on tasks where AI was allowed, while they believed AI had made them 20% faster. METR’s February 2026 follow-up gave estimates whose confidence intervals include zero, and the authors called the data very weak evidence. Measure your own team before you decide.",{"q":104,"a":105},"What is the break-even productivity gain for an AI coding setup?","Add the monthly tool cost, the review time other people spend on the AI output and the expected monthly cost of quality problems. Divide that sum by the engineer’s fully loaded monthly cost. The result is the minimum net gain, measured end to end, that the setup must deliver just to break even.",{"q":107,"a":108},"Can I send client code or personal data to a coding agent under the GDPR?","Only under a written contract that meets Article 28, with the vendor as processor. Check in writing that the vendor does not train on your content, what it retains and where processing happens. Some business plans, such as Claude Team, include no model training by default, but the contract still has to be in place.",{"q":110,"a":111},"Is one senior with agents cheaper than an agency?","It can be, but the answer depends on the measured gain, the agency’s day rate and the days the agency needs for the same scope. Put both options on the same scope and quality bar, and compare the cost per delivered day. I found no public study that compares the two directly.",[113,116,119,122,125,128,131,134,137],{"id":114,"title":115},"what-the-studies-measured","What the studies measured",{"id":117,"title":118},"surveys-trust-delivery","What the surveys say about trust and delivery",{"id":120,"title":121},"where-agents-pay-off","Where agents pay off, and where they do not",{"id":123,"title":124},"comparison-nobody-measured","The comparison nobody has measured",{"id":126,"title":127},"cost-model","A cost model you can fill in",{"id":129,"title":130},"risks","Three risks the spreadsheet will not show",{"id":132,"title":133},"when-one-senior-is-enough","When one senior with agents is enough",{"id":135,"title":136},"first-steps","What I would do first",{"id":138,"title":139},"sources","Sources",[141,145,148,151,155,158,161,164,167,170,173,176,177,180,183,186,270,271,274,299,327,328,331,334,335,338,347,405,408,411,414,417,424,425,428,431,434,442,445,448,456,457,460,467,468,483,486,487],{"type":142,"content":143},"paragraph",[144],"If you run engineering or finance, you will get this question soon: can one senior engineer with coding agents do the work of a team, or of an agency? Nobody has measured that comparison head to head. The studies that exist measure pieces of it, and some results surprised the people who ran them. My answer is a break-even test, not a verdict. One senior with agents wins only when the net productivity gain on your work exceeds the tool, review and quality costs, as a share of the engineer’s loaded cost.",{"type":146,"level":147,"id":114,"text":115},"heading",2,{"type":142,"content":149},[150],"The controlled trials with the largest gains measured code-completion assistants on bounded tasks. The 2025 trial measured an AI code editor with Claude models on real issues. Agents that plan, run commands and iterate on their own output have been studied less, so read each number as a snapshot of one tool, one year and one kind of task.",{"type":146,"level":152,"id":153,"text":154},3,"metr-trial","The trial that found a slowdown",{"type":142,"content":156},[157],"METR, a non-profit that evaluates AI systems, ran a randomised trial in early 2025 with 16 experienced open-source developers. They worked on 246 real issues in mature repositories, and each issue was randomly assigned to allow or forbid AI. Beforehand, the developers expected AI to cut their time by 24%. Afterwards they estimated a 20% cut. The measured time went the other way: tasks took 19% longer with AI, with a confidence interval from +2% to +39%.",{"type":142,"content":159},[160],"Two lessons follow. The developers were wrong about their own speed, so asking people how fast they feel is not a measurement. And the sample is small, the tools were early-2025 models and the code was familiar to the people working on it. The authors checked 20 properties of their setting and found that the slowdown held up across their analyses. They think AI may be useful elsewhere, for example for less experienced developers or in unfamiliar codebases.",{"type":146,"level":152,"id":162,"text":163},"metr-follow-up","The follow-up that settled nothing",{"type":142,"content":165},[166],"METR’s second experiment started in August 2025 with 57 developers, 143 repositories and more than 800 tasks. In February 2026 it called the data an unreliable signal. Developers who did not want to work without AI chose not to take part, which likely biases the estimate downwards. The pay also fell from $150 to $50 an hour, which may have changed who joined. The authors think developers are probably faster now, but their data is very weak evidence for the size of that gain, and their confidence intervals include zero.",{"type":146,"level":152,"id":168,"text":169},"controlled-gains","Controlled trials with large gains",{"type":142,"content":171},[172],"In a 2023 GitHub study, 95 developers were randomly split into a Copilot group and a control group, and everyone was asked to build an HTTP server in JavaScript as fast as they could. The Copilot group needed 71 minutes on average against 161 minutes, a 55.8% reduction with a 95% confidence interval from 21% to 89%. Only about 35 developers finished the task, and the success-rate difference was not statistically significant. Several authors work at GitHub or Microsoft Research, so the study is not independent of the vendor.",{"type":142,"content":174},[175],"The largest sample in this list comes from three companies. Researchers at Microsoft, Accenture and an anonymous Fortune 100 firm gave random subsets of developers an AI code-completion assistant during normal operations. Pooled across 4,867 developers, the authors estimate a 26.08% increase in completed tasks, with a standard error of 10.3%. Less experienced developers gained more, so a senior’s likely gain is below the average.",{"type":146,"level":147,"id":117,"text":118},{"type":142,"content":178},[179],"DORA’s 2024 report from Google Cloud is correlational, so read it as association, not cause. A 25% rise in AI adoption went with a 3.4% increase in code quality and a 3.1% increase in code review speed, but also with a 1.5% drop in delivery throughput and a 7.2% drop in delivery stability. 39% of respondents had little or no trust in AI-generated code. The 2025 report, based on nearly 5,000 technology professionals, makes a sharper point: AI amplifies what an organisation already does well or badly.",{"type":142,"content":181},[182],"The Stack Overflow Developer Survey 2025 shows the same tension in individuals. 84% of respondents use or plan to use AI tools, and 51% of professional developers use them daily. Yet 46% distrust the accuracy of AI tools, against 33% who trust them. The most common frustration, named by 66%, is output that is almost right but not quite, and 45% say debugging it takes more time.",{"type":142,"content":184},[185],"Table 1 sets the headline results beside their limits.",{"type":187,"head":188,"rows":197},"table",[189,191,193,195],[190],"Study",[192],"What it measured",[194],"Headline result",[196],"What it does not show",[198,207,216,225,234,243,252,261],[199,201,203,205],[200],"METR, 2025",[202],"16 experienced developers, 246 real issues",[204],"19% longer with AI",[206],"Late-2025 tools, other codebases",[208,210,212,214],[209],"METR, February 2026",[211],"57 developers, 800+ tasks",[213],"Confidence intervals include zero",[215],"A reliable speed-up figure",[217,219,221,223],[218],"GitHub, 2023 (Peng et al.)",[220],"95 randomised developers, one HTTP server task",[222],"55.8% less time",[224],"Maintenance work or long projects",[226,228,230,232],[227],"Microsoft, Accenture, Fortune 100 firm",[229],"4,867 developers in three field trials",[231],"+26.08% completed tasks",[233],"Agents, or quality after merge",[235,237,239,241],[236],"DORA, 2024",[238],"Survey of technology professionals",[240],"Per 25% more adoption: −1.5% throughput, −7.2% stability",[242],"Causation (correlational)",[244,246,248,250],[245],"Stack Overflow, 2025",[247],"Developers’ attitudes",[249],"46% distrust AI accuracy, 33% trust",[251],"Actual productivity",[253,255,257,259],[254],"Google, migrations, 2025",[256],"39 code migrations, 595 changes",[258],"74.45% of changes LLM-generated",[260],"Feature work; the half-time saving is an estimate",[262,264,266,268],[263],"Meta, TestGen-LLM, 2024",[265],"Unit tests for Instagram Reels and Stories",[267],"75% built, 57% passed reliably",[269],"Code outside test suites",{"type":146,"level":147,"id":120,"text":121},{"type":142,"content":272},[273],"The pattern is fairly consistent. Gains show up when the task is bounded, an automatic check decides whether the result is right, and a person only has to read the output. Gains shrink when the work depends on context outside the code, or when the codebase is large and familiar enough that every change needs a line-by-line check.",{"type":275,"ordered":276,"items":277},"list",false,[278,284,289,294],[279,283],{"tag":280,"children":281},"strong",[282],"Tests behind a filter."," Meta’s TestGen-LLM only proposes candidates that pass automated checks for a measurable improvement. On Instagram’s Reels and Stories, 75% of its test cases built, 57% passed reliably and 25% raised coverage. Engineers accepted 73% of its recommendations.",[285,288],{"tag":280,"children":286},[287],"Migrations."," In Google’s account of 39 migrations, 74.45% of the submitted changes and 69.46% of the edits were LLM-generated. Engineers estimated that total time fell by about half. That is an estimate, not a measurement.",[290,293],{"tag":280,"children":291},[292],"Bounded tasks with a clear finish line."," The 55.8% gain in the GitHub trial came from exactly that kind of task.",[295,298],{"tag":280,"children":296},[297],"Unfamiliar code and newer engineers."," The field trials found the largest gains among less experienced developers, and METR names unfamiliar codebases as a likely place for AI to help.",{"type":275,"ordered":276,"items":300},[301,306,311,322],[302,305],{"tag":280,"children":303},[304],"Familiar, mature code."," In METR’s trial, tasks in repositories the developers knew well took 19% longer with AI allowed.",[307,310],{"tag":280,"children":308},[309],"Unclear requirements."," The studies do not measure this. My reading is that the cost lands in review: the agent produces plausible code for the requirement it was given, and a senior has to check that it was the right one.",[312,315,316,321],{"tag":280,"children":313},[314],"Review and rework."," Almost-right output is the top frustration in the Stack Overflow survey. GitClear, which analyses code change data, reports that refactored lines fell from 25% of changed lines in 2021 to under 10% in 2024, while copy-pasted lines rose from 8.3% to 12.3%. That association is not proof of cause. For the review side, see my note on ",{"tag":317,"to":318,"children":319},"link","\u002Fblog\u002Fai-generated-pr-review-bottleneck",[320],"AI code review is the bottleneck now",".",[323,326],{"tag":280,"children":324},[325],"Knowledge outside the repository."," Pricing rules, regulatory logic and client quirks are not in the files an agent reads, and writing that context takes the senior’s time.",{"type":146,"level":147,"id":123,"text":124},{"type":142,"content":329},[330],"I found no study that compares one senior engineer with agents against a team or an agency on the same scope, quality bar and customer. The comparison has to be assembled from measured parts: the senior’s net gain, the costs the setup adds beyond the senior’s own time, and the price and pace of the alternative. The cost model below puts them in one place. It is a method for your numbers, not a forecast.",{"type":142,"content":332},[333],"The alternatives fail differently. A team costs more people, but more than one person knows the system and can review and ship, which a single senior does not provide. An agency sells capacity by the day, and its rate covers its own overhead and margin. Its days are only comparable when the quote includes the same work: discovery, tests, deployment and handover.",{"type":146,"level":147,"id":126,"text":127},{"type":142,"content":336},[337],"Use one scope, one period and one quality bar for every option. The diagram shows where the gain has to survive, and the table defines the inputs.",{"type":339,"attrs":340,"inner":344,"caption":345},"diagram",{"viewBox":341,"role":342,"aria-labelledby":343},"0 0 720 250","img","d2-gain-t d2-gain-d","\u003Ctitle id=\"d2-gain-t\">Where the gain has to survive\u003C\u002Ftitle>\u003Cdesc id=\"d2-gain-d\">Four stages in a row: draft, review, test and fix, and ship. Drafting is where the agent speeds work up, which is the gain. Review, testing and later incidents are costs. The net gain is read at the ship stage. Below the stages, the break-even rule says the net gain must exceed the tool, review and quality costs divided by the loaded cost.\u003C\u002Fdesc>\u003Ctext x=\"360\" y=\"30\" text-anchor=\"middle\" class=\"d-title\">Where the gain has to survive\u003C\u002Ftext>\u003Crect x=\"20\" y=\"60\" width=\"140\" height=\"70\" rx=\"10\" class=\"d-accent\" \u002F>\u003Ctext x=\"90\" y=\"90\" text-anchor=\"middle\" class=\"d-text\">Draft\u003C\u002Ftext>\u003Ctext x=\"90\" y=\"112\" text-anchor=\"middle\" class=\"d-small\">agent speeds this up\u003C\u002Ftext>\u003Crect x=\"205\" y=\"60\" width=\"140\" height=\"70\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"275\" y=\"90\" text-anchor=\"middle\" class=\"d-text\">Review\u003C\u002Ftext>\u003Ctext x=\"275\" y=\"112\" text-anchor=\"middle\" class=\"d-small\">a human reads it\u003C\u002Ftext>\u003Crect x=\"390\" y=\"60\" width=\"140\" height=\"70\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"460\" y=\"90\" text-anchor=\"middle\" class=\"d-text\">Test and fix\u003C\u002Ftext>\u003Ctext x=\"460\" y=\"112\" text-anchor=\"middle\" class=\"d-small\">bugs surface later\u003C\u002Ftext>\u003Crect x=\"575\" y=\"60\" width=\"140\" height=\"70\" rx=\"10\" class=\"d-mint\" \u002F>\u003Ctext x=\"645\" y=\"90\" text-anchor=\"middle\" class=\"d-text\">Ship\u003C\u002Ftext>\u003Ctext x=\"645\" y=\"112\" text-anchor=\"middle\" class=\"d-small\">measure it here\u003C\u002Ftext>\u003Cpath d=\"M162 95 H196\" class=\"d-line\" \u002F>\u003Cpath d=\"M205 95 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Cpath d=\"M347 95 H381\" class=\"d-line\" \u002F>\u003Cpath d=\"M390 95 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Cpath d=\"M532 95 H566\" class=\"d-line\" \u002F>\u003Cpath d=\"M575 95 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Ctext x=\"90\" y=\"156\" text-anchor=\"middle\" class=\"d-label\">gain\u003C\u002Ftext>\u003Ctext x=\"275\" y=\"156\" text-anchor=\"middle\" class=\"d-label\">cost V\u003C\u002Ftext>\u003Ctext x=\"460\" y=\"156\" text-anchor=\"middle\" class=\"d-label\">cost Q\u003C\u002Ftext>\u003Ctext x=\"645\" y=\"156\" text-anchor=\"middle\" class=\"d-label\">net gain g\u003C\u002Ftext>\u003Crect x=\"20\" y=\"186\" width=\"680\" height=\"44\" rx=\"10\" class=\"d-fill-muted\" \u002F>\u003Ctext x=\"360\" y=\"213\" text-anchor=\"middle\" class=\"d-text\">Break-even: g must exceed (T + V + Q) \u002F L\u003C\u002Ftext>",[346],"The gain is made at the draft and spent at review, testing and incidents. Measure it from ticket to production.",{"type":187,"head":348,"rows":355},[349,351,353],[350],"Input",[352],"What to enter",[354],"Where it comes from",[356,363,370,377,384,391,398],[357,359,361],[358],"L, loaded cost",[360],"Monthly cost of the senior: salary, employer costs, equipment, overhead",[362],"Payroll and finance",[364,366,368],[365],"T, tool cost",[367],"Seats, usage above the plan, API spend, hosting for agents",[369],"Vendor invoices",[371,373,375],[372],"V, review cost",[374],"Hours others spend reviewing and fixing agent output, times their hourly cost",[376],"Pull request time logs",[378,380,382],[379],"Q, quality cost",[381],"Expected monthly cost of incidents, rework and customer credits",[383],"Incident and defect log",[385,387,389],[386],"B, baseline",[388],"Days of scoped work delivered per month before agents",[390],"Three months of tracking",[392,394,396],[393],"g, net gain",[395],"Measured gain from ticket to production at the same quality, as a fraction: 0.10 is 10%",[397],"A pilot, not a survey",[399,401,403],[400],"D and Sa, agency",[402],"Agency day rate, and the days it quotes for the same scope",[404],"Written quote",{"type":406,"code":407},"code","cost per scope day, no agents      = L \u002F B\ncost per scope day, with agents    = (L + T + V + Q) \u002F (B × (1 + g))\nbreak-even net gain                = (T + V + Q) \u002F L\nsenior with agents, scope S days   = (L + T + V + Q) × S \u002F (B × (1 + g))\nagency, same scope                 = D × Sa",{"type":142,"content":409},[410],"The break-even line is the number to take into the room. Every point of tool, review and quality cost, as a share of the loaded cost, must be earned back as a point of measured net gain. If those costs add up to 10% of the loaded cost, a net gain below 10% makes each delivered day more expensive than before. For a team, add the loaded costs and use the team’s measured output.",{"type":142,"content":412},[413],"The tool line is the easiest to estimate and the least stable. As of October 2026, Claude Pro costs $17 a month on an annual plan, or $20 billed monthly. A Claude Team standard seat costs $20 a month billed annually, a premium seat $100 a month billed annually, and Claude Max starts at $100 a month. GitHub Copilot Pro costs $10 a month, Pro+ $39 and Max $100. Usage limits apply, and Anthropic says its prices and plans may change at its discretion.",{"type":142,"content":415},[416],"Compare the options on the same scope. The agency side is its day rate times its quoted days; the senior side is the scope formula. The answer flips in one of two ways: the measured gain is large and review cost is low, or the agency quotes far more days than the work needs. Ask both sides for the same deliverables before comparing numbers.",{"type":418,"variant":419,"title":420,"body":421},"callout","tip","Same scope, same bar",[422],[423],"Put the pilot’s measured gain in the sheet, not a vendor’s benchmark, and re-run it whenever the tool or the model changes.",{"type":146,"level":147,"id":129,"text":130},{"type":146,"level":152,"id":426,"text":427},"single-point-of-failure","Single point of failure",{"type":142,"content":429},[430],"One senior is a bus factor of one, and agents deepen that dependency, because the know-how now sits in prompts, skills and configuration as well as in one head. Keep the agent configuration, the specs and the review rules in the repository. Name a second person who can review and ship, and agree in advance what happens during illness or holidays. The vendor is a second single point, so price a fallback, such as an agency retainer, and do not assume today’s price.",{"type":146,"level":152,"id":432,"text":433},"quality-debt","Quality debt",{"type":142,"content":435},[436,437,441],"Quality debt arrives later and does not show up in coding time. DORA’s association with lower stability is the warning: code can arrive faster than the system absorbs it. The defence belongs in the pipeline, not the prompt. Make the merge depend on checks the agent cannot edit (my note on ",{"tag":317,"to":438,"children":439},"\u002Fblog\u002Fharness-engineering-coding-agents",[440],"harness engineering for coding agents"," covers the set-up), require a human approval for every change, and track rework, such as changes reverted or reopened within a fixed window.",{"type":146,"level":152,"id":443,"text":444},"data-protection","Data protection",{"type":142,"content":446},[447],"Article 28 of the GDPR applies when a vendor processes personal data for you. The processor needs a written contract that sets out the subject-matter and duration of the processing, its nature and purpose, and the type of personal data. It may not bring in another processor without your prior written authorisation, specific or general. Put the agent vendor under that contract before personal data reaches a prompt, including names in tickets, customer records in test fixtures and personal data in logs.",{"type":142,"content":449},[450,451,455],"Vendor terms matter too. Anthropic’s commercial terms say it may not train models on Customer Content from its services, and its Team plan lists no model training on your content by default. Customers in the EEA, Switzerland or the UK contract with Anthropic Ireland. Anthropic’s data-residency documentation describes US-only inference at 1.1 times the standard price, with global routing at standard pricing otherwise. In the page I read I found no EU-only option, so ask for one in writing. My note on ",{"tag":317,"to":452,"children":453},"\u002Fblog\u002Fgdpr-llm-api-eu-data-residency",[454],"GDPR and LLM API data residency"," covers the EU options in more detail.",{"type":146,"level":147,"id":132,"text":133},{"type":142,"content":458},[459],"My rule of thumb is the chain below. Work through it in order and stop at the first no. The first three questions decide whether the setup is safe to run; the last one decides whether it is cheaper.",{"type":339,"attrs":461,"inner":464,"caption":465},{"viewBox":462,"role":342,"aria-labelledby":463},"0 0 720 390","d1-decide-t d1-decide-d","\u003Ctitle id=\"d1-decide-t\">Deciding whether one senior with agents is enough\u003C\u002Ftitle>\u003Cdesc id=\"d1-decide-d\">A chain of five steps, read from the top. If the scope is not written down, write the spec first. If tests do not catch regressions, fix the tests and CI first. If one person cannot review and ship, add a reviewer or a retainer. If the net gain is not above break-even, keep the team or hire the agency. If every check passes, take the work and re-measure every quarter.\u003C\u002Fdesc>\u003Ctext x=\"700\" y=\"22\" text-anchor=\"end\" class=\"d-label\">Start at the top, stop at the first no\u003C\u002Ftext>\u003Crect x=\"20\" y=\"40\" width=\"300\" height=\"46\" rx=\"10\" class=\"d-gold\" \u002F>\u003Ctext x=\"170\" y=\"68\" text-anchor=\"middle\" class=\"d-text\">Is the scope written down?\u003C\u002Ftext>\u003Cpath d=\"M320 63 H392\" class=\"d-line-accent\" \u002F>\u003Cpath d=\"M400 63 l-9 -5 v10 z\" class=\"d-head-accent\" \u002F>\u003Ctext x=\"332\" y=\"57\" class=\"d-label\">no\u003C\u002Ftext>\u003Crect x=\"400\" y=\"40\" width=\"300\" height=\"46\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"550\" y=\"59\" text-anchor=\"middle\" class=\"d-text\">Write the spec first\u003C\u002Ftext>\u003Ctext x=\"550\" y=\"76\" text-anchor=\"middle\" class=\"d-small\">A product owner or the team writes it\u003C\u002Ftext>\u003Cpath d=\"M170 86 V103\" class=\"d-line\" \u002F>\u003Cpath d=\"M170 112 l-5 -9 h10 z\" class=\"d-head\" \u002F>\u003Ctext x=\"180\" y=\"98\" class=\"d-label\">yes\u003C\u002Ftext>\u003Crect x=\"20\" y=\"112\" width=\"300\" height=\"46\" rx=\"10\" class=\"d-gold\" \u002F>\u003Ctext x=\"170\" y=\"140\" text-anchor=\"middle\" class=\"d-text\">Do tests catch regressions?\u003C\u002Ftext>\u003Cpath d=\"M320 135 H392\" class=\"d-line-accent\" \u002F>\u003Cpath d=\"M400 135 l-9 -5 v10 z\" class=\"d-head-accent\" \u002F>\u003Ctext x=\"332\" y=\"129\" class=\"d-label\">no\u003C\u002Ftext>\u003Crect x=\"400\" y=\"112\" width=\"300\" height=\"46\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"550\" y=\"131\" text-anchor=\"middle\" class=\"d-text\">Fix the tests and CI first\u003C\u002Ftext>\u003Ctext x=\"550\" y=\"148\" text-anchor=\"middle\" class=\"d-small\">A failing test must stop the merge\u003C\u002Ftext>\u003Cpath d=\"M170 158 V175\" class=\"d-line\" \u002F>\u003Cpath d=\"M170 184 l-5 -9 h10 z\" class=\"d-head\" \u002F>\u003Ctext x=\"180\" y=\"170\" class=\"d-label\">yes\u003C\u002Ftext>\u003Crect x=\"20\" y=\"184\" width=\"300\" height=\"46\" rx=\"10\" class=\"d-gold\" \u002F>\u003Ctext x=\"170\" y=\"212\" text-anchor=\"middle\" class=\"d-text\">Can a second person review and ship?\u003C\u002Ftext>\u003Cpath d=\"M320 207 H392\" class=\"d-line-accent\" \u002F>\u003Cpath d=\"M400 207 l-9 -5 v10 z\" class=\"d-head-accent\" \u002F>\u003Ctext x=\"332\" y=\"201\" class=\"d-label\">no\u003C\u002Ftext>\u003Crect x=\"400\" y=\"184\" width=\"300\" height=\"46\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"550\" y=\"203\" text-anchor=\"middle\" class=\"d-text\">Add a reviewer or a retainer\u003C\u002Ftext>\u003Ctext x=\"550\" y=\"220\" text-anchor=\"middle\" class=\"d-small\">No single point of failure\u003C\u002Ftext>\u003Cpath d=\"M170 230 V247\" class=\"d-line\" \u002F>\u003Cpath d=\"M170 256 l-5 -9 h10 z\" class=\"d-head\" \u002F>\u003Ctext x=\"180\" y=\"242\" class=\"d-label\">yes\u003C\u002Ftext>\u003Crect x=\"20\" y=\"256\" width=\"300\" height=\"46\" rx=\"10\" class=\"d-gold\" \u002F>\u003Ctext x=\"170\" y=\"284\" text-anchor=\"middle\" class=\"d-text\">Is the net gain above break-even?\u003C\u002Ftext>\u003Cpath d=\"M320 279 H392\" class=\"d-line-accent\" \u002F>\u003Cpath d=\"M400 279 l-9 -5 v10 z\" class=\"d-head-accent\" \u002F>\u003Ctext x=\"332\" y=\"273\" class=\"d-label\">no\u003C\u002Ftext>\u003Crect x=\"400\" y=\"256\" width=\"300\" height=\"46\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"550\" y=\"275\" text-anchor=\"middle\" class=\"d-text\">Keep the team or hire the agency\u003C\u002Ftext>\u003Ctext x=\"550\" y=\"292\" text-anchor=\"middle\" class=\"d-small\">Compare the same scope and bar\u003C\u002Ftext>\u003Cpath d=\"M170 302 V319\" class=\"d-line\" \u002F>\u003Cpath d=\"M170 328 l-5 -9 h10 z\" class=\"d-head\" \u002F>\u003Ctext x=\"180\" y=\"314\" class=\"d-label\">yes\u003C\u002Ftext>\u003Crect x=\"20\" y=\"328\" width=\"300\" height=\"46\" rx=\"10\" class=\"d-mint\" \u002F>\u003Ctext x=\"170\" y=\"356\" text-anchor=\"middle\" class=\"d-text\">Take it and re-measure quarterly\u003C\u002Ftext>",[466],"Stop at the first no. An unsafe setup is not cheaper, whatever the measured gain.",{"type":146,"level":147,"id":135,"text":136},{"type":275,"ordered":469,"items":470},true,[471,473,475,477,479,481],[472],"Write down the baseline. For three months, record the days of scoped work this person delivers and the time from ticket to production.",[474],"Run a two-week pilot on one bounded project with agents, and measure the end-to-end gain against the baseline, not the feeling of speed.",[476],"Fill in the cost model with real numbers, and compare it with a written quote from an agency for the same scope.",[478],"Sign the data processing agreement, and confirm training, retention and processing region in writing before client data goes into any prompt.",[480],"Name a second person who reviews and ships agent work, and write down what happens when the senior is away.",[482],"Re-measure every quarter. Tools change faster than the studies, and METR’s own follow-up shows how hard the number is to pin down.",{"type":142,"content":484},[485],"None of this needs a platform. It needs a baseline, a measured gain and a signed contract, in that order.",{"type":146,"level":147,"id":138,"text":139},{"type":275,"ordered":469,"items":488},[489,493,496,499,502,505,508,511,514,517,520,523,526,529,532,535],[490],{"tag":491,"href":36,"children":492},"a",[35],[494],{"tag":491,"href":39,"children":495},[38],[497],{"tag":491,"href":42,"children":498},[41],[500],{"tag":491,"href":45,"children":501},[44],[503],{"tag":491,"href":48,"children":504},[47],[506],{"tag":491,"href":51,"children":507},[50],[509],{"tag":491,"href":54,"children":510},[53],[512],{"tag":491,"href":57,"children":513},[56],[515],{"tag":491,"href":60,"children":516},[59],[518],{"tag":491,"href":63,"children":519},[62],[521],{"tag":491,"href":66,"children":522},[65],[524],{"tag":491,"href":69,"children":525},[68],[527],{"tag":491,"href":72,"children":528},[71],[530],{"tag":491,"href":75,"children":531},[74],[533],{"tag":491,"href":78,"children":534},[77],[536],{"tag":491,"href":81,"children":537},[80],[539,616,671,750],{"slug":540,"published":541,"minutes":6,"category":7,"tags":542,"keywords":548,"about":557,"sources":569,"cover":610,"og":611,"expertise":84,"locales":612,"lang":86,"title":613,"description":614,"coverAlt":615},"spec-driven-development-coding-agents","2026-09-30",[543,544,545,546,547],"Spec-driven development","Coding agents","Acceptance criteria","Plan mode","AI engineering",[549,550,551,552,553,554,555,556],"spec-driven development","spec driven development coding agents","acceptance criteria for coding agents","AI coding agent plan mode","GitHub Spec Kit","Kiro specs requirements design tasks","vibe coding vs spec-driven development","EARS requirements syntax",[558,561,563,566],{"name":559,"url":560},"Vibe coding","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FVibe_coding",{"name":553,"url":562},"https:\u002F\u002Fgithub.com\u002Fgithub\u002Fspec-kit",{"name":564,"url":565},"EARS (Easy Approach to Requirements Syntax)","https:\u002F\u002Falistairmavin.com\u002Fears\u002F",{"name":567,"url":568},"Kiro","https:\u002F\u002Fkiro.dev",[570,573,575,578,581,584,586,589,592,595,597,599,602,604,607],{"title":571,"url":572},"Birgitta Böckeler: Understanding Spec-Driven-Development: Kiro, spec-kit, and Tessl (martinfowler.com, 15 October 2025)","https:\u002F\u002Fmartinfowler.com\u002Farticles\u002Fexploring-gen-ai\u002Fsdd-3-tools.html",{"title":574,"url":562},"GitHub Spec Kit: README",{"title":576,"url":577},"GitHub Spec Kit: spec-driven.md methodology","https:\u002F\u002Fgithub.com\u002Fgithub\u002Fspec-kit\u002Fblob\u002Fmain\u002Fspec-driven.md",{"title":579,"url":580},"Kiro documentation: specs","https:\u002F\u002Fkiro.dev\u002Fdocs\u002Fspecs\u002F",{"title":582,"url":583},"Kiro: Introducing Kiro","https:\u002F\u002Fkiro.dev\u002Fblog\u002Fintroducing-kiro\u002F",{"title":585,"url":565},"Alistair Mavin: EARS, the Easy Approach to Requirements Syntax",{"title":587,"url":588},"Anthropic: Best practices for Claude Code","https:\u002F\u002Fcode.claude.com\u002Fdocs\u002Fen\u002Fbest-practices",{"title":590,"url":591},"Cline documentation: Plan and Act","https:\u002F\u002Fdocs.cline.bot\u002Ffeatures\u002Fplan-and-act",{"title":593,"url":594},"OpenAI: Codex slash command reference","https:\u002F\u002Flearn.chatgpt.com\u002Fdocs\u002Freference\u002Fslash-commands",{"title":596,"url":36},"METR: early-2025 AI and experienced open-source developer productivity (July 2025)",{"title":598,"url":42},"METR: We are changing our developer productivity experiment design (February 2026)",{"title":600,"url":601},"Thoughtworks Technology Radar: Spec-driven development","https:\u002F\u002Fwww.thoughtworks.com\u002Fradar\u002Ftechniques\u002Fspec-driven-development",{"title":603,"url":560},"Wikipedia: Vibe coding",{"title":605,"url":606},"Anthropic: Claude API pricing","https:\u002F\u002Fplatform.claude.com\u002Fdocs\u002Fen\u002Fabout-claude\u002Fpricing",{"title":608,"url":609},"Anthropic: Prompt caching","https:\u002F\u002Fplatform.claude.com\u002Fdocs\u002Fen\u002Fbuild-with-claude\u002Fprompt-caching","\u002Fimages\u002Fblog\u002Fspec-driven-development-coding-agents\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fspec-driven-development-coding-agents\u002Fog.jpg",[86,87,88],"Spec-driven development for coding agents: agree the plan before the code","Vibe coding breaks on real codebases. Write a spec with acceptance criteria, a plan and tasks, let the agent tick them off, and review before the first line of code.","Cover art for spec-driven development: a pipeline from spec and plan to tasks and verification, with a review gate before any code is written.",{"slug":617,"published":618,"minutes":619,"category":7,"tags":620,"keywords":626,"about":636,"sources":646,"cover":665,"og":666,"expertise":84,"locales":667,"lang":86,"title":668,"description":669,"coverAlt":670},"mcp-tool-design-lessons-jira-server","2026-09-18",10,[621,622,623,624,625],"MCP","Tool design","Context engineering","Jira","Agents",[627,628,629,630,631,632,633,634,635],"MCP tool design","MCP best practices","MCP tool descriptions","agent tool selection","MCP context bloat","how many tools should an MCP server have","MCP tool definition token cost","MCP error handling isError","Jira MCP server",[637,640,643],{"name":638,"url":639},"Model Context Protocol","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FModel_Context_Protocol",{"name":641,"url":642},"Jira (software)","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FJira_(software)",{"name":644,"url":645},"Intelligent agent","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FIntelligent_agent",[647,650,653,656,659,662],{"title":648,"url":649},"Writing effective tools for agents – with agents (Anthropic)","https:\u002F\u002Fwww.anthropic.com\u002Fengineering\u002Fwriting-tools-for-agents",{"title":651,"url":652},"Introducing advanced tool use on the Claude Developer Platform (Anthropic)","https:\u002F\u002Fwww.anthropic.com\u002Fengineering\u002Fadvanced-tool-use",{"title":654,"url":655},"Code execution with MCP (Anthropic)","https:\u002F\u002Fwww.anthropic.com\u002Fengineering\u002Fcode-execution-with-mcp",{"title":657,"url":658},"MCP vs CLI: context window cost (Blocks.ai)","https:\u002F\u002Fblocks.ai\u002Fblog\u002Fmcp-vs-cli-context-window-cost",{"title":660,"url":661},"Demystifying evals for AI agents (Anthropic)","https:\u002F\u002Fwww.anthropic.com\u002Fengineering\u002Fdemystifying-evals-for-ai-agents",{"title":663,"url":664},"MCP 2026-07-28 specification: Tools","https:\u002F\u002Fmodelcontextprotocol.io\u002Fspecification\u002F2026-07-28\u002Fserver\u002Ftools","\u002Fimages\u002Fblog\u002Fmcp-tool-design-lessons-jira-server\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fmcp-tool-design-lessons-jira-server\u002Fog.jpg",[86,87,88],"MCP tool design: lessons from a 20-tool Jira server","MCP tool design that agents get right: token cost of tool definitions, when to merge tools, naming, concise output, errors that steer and a small selection eval.","Network diagram with a Jira MCP server at the hub and five satellites: search, create, transition, comments and test runs",{"slug":672,"published":673,"minutes":674,"category":7,"tags":675,"keywords":678,"about":689,"sources":695,"cover":744,"og":745,"expertise":84,"locales":746,"lang":86,"title":747,"description":748,"coverAlt":749},"ai-agent-memory-design","2026-09-10",13,[676,623,677,13],"AI agent memory","Memory poisoning",[679,680,681,682,683,684,685,686,687,688],"AI agent memory design","long-term memory for AI agents","episodic semantic procedural memory LLM","agent memory architecture","ChatGPT memory vs Claude memory","Claude memory tool","AI memory poisoning","LLM context compaction","AI agent memory GDPR","short-term vs long-term memory agents",[690,691,692],{"name":644,"url":645},{"name":25,"url":26},{"name":693,"url":694},"Prompt injection","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FPrompt_injection",[696,699,702,705,708,711,714,717,720,723,726,729,732,735,738,741],{"title":697,"url":698},"Anthropic docs: Memory tool","https:\u002F\u002Fplatform.claude.com\u002Fdocs\u002Fen\u002Fagents-and-tools\u002Ftool-use\u002Fmemory-tool",{"title":700,"url":701},"Anthropic docs: Context editing","https:\u002F\u002Fplatform.claude.com\u002Fdocs\u002Fen\u002Fbuild-with-claude\u002Fcontext-editing",{"title":703,"url":704},"Anthropic Engineering: Effective context engineering for AI agents","https:\u002F\u002Fwww.anthropic.com\u002Fengineering\u002Feffective-context-engineering-for-ai-agents",{"title":706,"url":707},"Sumers et al.: Cognitive Architectures for Language Agents (CoALA)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2309.02427",{"title":709,"url":710},"Packer et al.: MemGPT, Towards LLMs as Operating Systems","https:\u002F\u002Farxiv.org\u002Fabs\u002F2310.08560",{"title":712,"url":713},"Park et al.: Generative Agents, Interactive Simulacra of Human Behavior","https:\u002F\u002Farxiv.org\u002Fabs\u002F2304.03442",{"title":715,"url":716},"Unit 42: When AI Remembers Too Much, persistent behaviors in agents memory","https:\u002F\u002Funit42.paloaltonetworks.com\u002Findirect-prompt-injection-poisons-ai-longterm-memory\u002F",{"title":718,"url":719},"From Untrusted Input to Trusted Memory: A Systematic Study of Memory Poisoning Attacks in LLM Agents (preprint)","https:\u002F\u002Farxiv.org\u002Fhtml\u002F2606.04329v1",{"title":721,"url":722},"The Hacker News: ChatGPT macOS flaw could have enabled long-term spyware via memory function","https:\u002F\u002Fthehackernews.com\u002F2024\u002F09\u002Fchatgpt-macos-flaw-couldve-enabled-long.html",{"title":724,"url":725},"Vectorize: OWASP ASI06, Memory and Context Poisoning explained","https:\u002F\u002Fvectorize.io\u002Farticles\u002Fowasp-asi06",{"title":727,"url":728},"Claude Help Center: Use chat search and memory to build on previous context","https:\u002F\u002Fsupport.claude.com\u002Fen\u002Farticles\u002F11817273-use-claude-s-chat-search-and-memory-to-build-on-previous-context",{"title":730,"url":731},"OpenAI Help Center: Memory in ChatGPT","https:\u002F\u002Fhelp.openai.com\u002Fen\u002Farticles\u002F8590148-memory-faq",{"title":733,"url":734},"OpenAI Help Center: Dots privacy, security, and safety FAQs","https:\u002F\u002Fhelp.openai.com\u002Fen\u002Farticles\u002F20001529-dots-privacy-security-and-safety-faqs",{"title":736,"url":737},"Flavio Copes: A deep dive into OpenAI dots (quotes the dots documentation on memory)","https:\u002F\u002Fflaviocopes.com\u002Fopenai-dots\u002F",{"title":739,"url":740},"GDPR Article 5: Principles relating to processing of personal data","https:\u002F\u002Fgdpr-info.eu\u002Fart-5-gdpr\u002F",{"title":742,"url":743},"GDPR Article 17: Right to erasure","https:\u002F\u002Fgdpr-info.eu\u002Fart-17-gdpr\u002F","\u002Fimages\u002Fblog\u002Fai-agent-memory-design\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fai-agent-memory-design\u002Fog.jpg",[86,87,88],"Designing memory for AI agents: tiers, write rules, poisoning and GDPR","How to design AI agent memory: context vs session vs long-term tiers, what to write and never store, retrieval, compaction, poisoning and GDPR erasure.","Diagram: nested memory layers of an AI agent, from the working context window through session state to long-term episodic and semantic memory.",{"slug":751,"published":752,"minutes":753,"category":7,"tags":754,"keywords":759,"about":769,"sources":778,"cover":803,"og":804,"expertise":84,"locales":805,"lang":86,"title":806,"description":807,"coverAlt":808},"harness-engineering-coding-agents","2026-09-04",8,[755,544,756,757,758],"Harness engineering","Code quality","Mutation testing","TDD",[760,761,762,763,764,765,766,767,768],"harness engineering","harness engineering coding agents","AI code quality","coding agent guardrails","mutation testing AI generated tests","red\u002Fgreen TDD with AI agents","guides and sensors coding agents","how to make AI agent pull requests mergeable","can I trust tests written by AI",[770,772,775],{"name":757,"url":771},"https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FMutation_testing",{"name":773,"url":774},"Test-driven development","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FTest-driven_development",{"name":776,"url":777},"Static program analysis","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FStatic_program_analysis",[779,782,785,788,791,794,797,800],{"title":780,"url":781},"Birgitta Böckeler: Harness engineering for coding agent users (Apr 2026)","https:\u002F\u002Fmartinfowler.com\u002Farticles\u002Fharness-engineering.html",{"title":783,"url":784},"Birgitta Böckeler: Maintainability sensors for coding agents (May 2026)","https:\u002F\u002Fmartinfowler.com\u002Farticles\u002Fsensors-for-coding-agents.html",{"title":786,"url":787},"Simon Willison: Agentic Engineering Patterns","https:\u002F\u002Fsimonwillison.net\u002Fguides\u002Fagentic-engineering-patterns\u002F",{"title":789,"url":790},"Simon Willison: First run the tests","https:\u002F\u002Fsimonwillison.net\u002Fguides\u002Fagentic-engineering-patterns\u002Ffirst-run-the-tests\u002F",{"title":792,"url":793},"Anthropic: Effective harnesses for long-running agents (Nov 2025)","https:\u002F\u002Fwww.anthropic.com\u002Fengineering\u002Feffective-harnesses-for-long-running-agents",{"title":795,"url":796},"Claude Code docs: How Claude remembers your project","https:\u002F\u002Fcode.claude.com\u002Fdocs\u002Fen\u002Fmemory",{"title":798,"url":799},"Stryker Mutator documentation","https:\u002F\u002Fstryker-mutator.io\u002Fdocs\u002F",{"title":801,"url":802},"Infection: command line options","https:\u002F\u002Finfection.github.io\u002Fguide\u002Fcommand-line-options.html","\u002Fimages\u002Fblog\u002Fharness-engineering-coding-agents\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fharness-engineering-coding-agents\u002Fog.jpg",[86,87,88],"Harness engineering: guides and sensors that make agent PRs mergeable","Harness engineering for coding agents: guides and sensors, where to run each check, red\u002Fgreen TDD, and mutation testing to verify the tests the agent wrote.","Concentric rings around a coding agent's model: behaviour, architecture fitness and maintainability harnesses, from outside in.",1791636875215]