Blog/AI agents

One senior with coding agents versus a team: what the evidence says

The METR, DORA, Microsoft and GitHub studies on AI coding tools, what they do not prove, and a break-even cost model for one senior versus a team or agency.

··11 min read

  • AI coding agents
  • Developer productivity
  • Engineering economics
  • Team cost
  • GDPR
Cover art for AI coding economics: a draft-to-ship pipeline where agents speed up drafting, and review and quality costs take part of the gain back.

Key takeaways

  • The studies disagree, and the pattern is useful: in a 2025 trial, experienced developers on codebases they knew took 19% longer with AI, while in a 2023 trial a single bounded task took 55.8% less time with Copilot.
  • Measure the net gain end to end, from ticket to production at the same quality. Speed gained at the draft can be spent again in review, rework and incidents.
  • The break-even is simple: the net gain has to beat the tool, review and quality costs, expressed as a share of the engineer’s loaded cost.
  • One senior with agents concentrates risk in one person and one vendor. Price in a second reviewer, a written runbook and a fallback before you decide.
  • Sign a data processing agreement and confirm training and processing region in writing before client data goes into a prompt, as Article 28 of the GDPR requires for processors.

Listen to this article

0:000:00

If you run engineering or finance, you will get this question soon: can one senior engineer with coding agents do the work of a team, or of an agency? Nobody has measured that comparison head to head. The studies that exist measure pieces of it, and some results surprised the people who ran them. My answer is a break-even test, not a verdict. One senior with agents wins only when the net productivity gain on your work exceeds the tool, review and quality costs, as a share of the engineer’s loaded cost.

What the studies measured

The controlled trials with the largest gains measured code-completion assistants on bounded tasks. The 2025 trial measured an AI code editor with Claude models on real issues. Agents that plan, run commands and iterate on their own output have been studied less, so read each number as a snapshot of one tool, one year and one kind of task.

The trial that found a slowdown

METR, a non-profit that evaluates AI systems, ran a randomised trial in early 2025 with 16 experienced open-source developers. They worked on 246 real issues in mature repositories, and each issue was randomly assigned to allow or forbid AI. Beforehand, the developers expected AI to cut their time by 24%. Afterwards they estimated a 20% cut. The measured time went the other way: tasks took 19% longer with AI, with a confidence interval from +2% to +39%.

Two lessons follow. The developers were wrong about their own speed, so asking people how fast they feel is not a measurement. And the sample is small, the tools were early-2025 models and the code was familiar to the people working on it. The authors checked 20 properties of their setting and found that the slowdown held up across their analyses. They think AI may be useful elsewhere, for example for less experienced developers or in unfamiliar codebases.

The follow-up that settled nothing

METR’s second experiment started in August 2025 with 57 developers, 143 repositories and more than 800 tasks. In February 2026 it called the data an unreliable signal. Developers who did not want to work without AI chose not to take part, which likely biases the estimate downwards. The pay also fell from $150 to $50 an hour, which may have changed who joined. The authors think developers are probably faster now, but their data is very weak evidence for the size of that gain, and their confidence intervals include zero.

Controlled trials with large gains

In a 2023 GitHub study, 95 developers were randomly split into a Copilot group and a control group, and everyone was asked to build an HTTP server in JavaScript as fast as they could. The Copilot group needed 71 minutes on average against 161 minutes, a 55.8% reduction with a 95% confidence interval from 21% to 89%. Only about 35 developers finished the task, and the success-rate difference was not statistically significant. Several authors work at GitHub or Microsoft Research, so the study is not independent of the vendor.

The largest sample in this list comes from three companies. Researchers at Microsoft, Accenture and an anonymous Fortune 100 firm gave random subsets of developers an AI code-completion assistant during normal operations. Pooled across 4,867 developers, the authors estimate a 26.08% increase in completed tasks, with a standard error of 10.3%. Less experienced developers gained more, so a senior’s likely gain is below the average.

What the surveys say about trust and delivery

DORA’s 2024 report from Google Cloud is correlational, so read it as association, not cause. A 25% rise in AI adoption went with a 3.4% increase in code quality and a 3.1% increase in code review speed, but also with a 1.5% drop in delivery throughput and a 7.2% drop in delivery stability. 39% of respondents had little or no trust in AI-generated code. The 2025 report, based on nearly 5,000 technology professionals, makes a sharper point: AI amplifies what an organisation already does well or badly.

The Stack Overflow Developer Survey 2025 shows the same tension in individuals. 84% of respondents use or plan to use AI tools, and 51% of professional developers use them daily. Yet 46% distrust the accuracy of AI tools, against 33% who trust them. The most common frustration, named by 66%, is output that is almost right but not quite, and 45% say debugging it takes more time.

Table 1 sets the headline results beside their limits.

StudyWhat it measuredHeadline resultWhat it does not show
METR, 202516 experienced developers, 246 real issues19% longer with AILate-2025 tools, other codebases
METR, February 202657 developers, 800+ tasksConfidence intervals include zeroA reliable speed-up figure
GitHub, 2023 (Peng et al.)95 randomised developers, one HTTP server task55.8% less timeMaintenance work or long projects
Microsoft, Accenture, Fortune 100 firm4,867 developers in three field trials+26.08% completed tasksAgents, or quality after merge
DORA, 2024Survey of technology professionalsPer 25% more adoption: −1.5% throughput, −7.2% stabilityCausation (correlational)
Stack Overflow, 2025Developers’ attitudes46% distrust AI accuracy, 33% trustActual productivity
Google, migrations, 202539 code migrations, 595 changes74.45% of changes LLM-generatedFeature work; the half-time saving is an estimate
Meta, TestGen-LLM, 2024Unit tests for Instagram Reels and Stories75% built, 57% passed reliablyCode outside test suites

Where agents pay off, and where they do not

The pattern is fairly consistent. Gains show up when the task is bounded, an automatic check decides whether the result is right, and a person only has to read the output. Gains shrink when the work depends on context outside the code, or when the codebase is large and familiar enough that every change needs a line-by-line check.

  • Tests behind a filter. Meta’s TestGen-LLM only proposes candidates that pass automated checks for a measurable improvement. On Instagram’s Reels and Stories, 75% of its test cases built, 57% passed reliably and 25% raised coverage. Engineers accepted 73% of its recommendations.
  • Migrations. In Google’s account of 39 migrations, 74.45% of the submitted changes and 69.46% of the edits were LLM-generated. Engineers estimated that total time fell by about half. That is an estimate, not a measurement.
  • Bounded tasks with a clear finish line. The 55.8% gain in the GitHub trial came from exactly that kind of task.
  • Unfamiliar code and newer engineers. The field trials found the largest gains among less experienced developers, and METR names unfamiliar codebases as a likely place for AI to help.
  • Familiar, mature code. In METR’s trial, tasks in repositories the developers knew well took 19% longer with AI allowed.
  • Unclear requirements. The studies do not measure this. My reading is that the cost lands in review: the agent produces plausible code for the requirement it was given, and a senior has to check that it was the right one.
  • Review and rework. Almost-right output is the top frustration in the Stack Overflow survey. GitClear, which analyses code change data, reports that refactored lines fell from 25% of changed lines in 2021 to under 10% in 2024, while copy-pasted lines rose from 8.3% to 12.3%. That association is not proof of cause. For the review side, see my note on AI code review is the bottleneck now.
  • Knowledge outside the repository. Pricing rules, regulatory logic and client quirks are not in the files an agent reads, and writing that context takes the senior’s time.

The comparison nobody has measured

I found no study that compares one senior engineer with agents against a team or an agency on the same scope, quality bar and customer. The comparison has to be assembled from measured parts: the senior’s net gain, the costs the setup adds beyond the senior’s own time, and the price and pace of the alternative. The cost model below puts them in one place. It is a method for your numbers, not a forecast.

The alternatives fail differently. A team costs more people, but more than one person knows the system and can review and ship, which a single senior does not provide. An agency sells capacity by the day, and its rate covers its own overhead and margin. Its days are only comparable when the quote includes the same work: discovery, tests, deployment and handover.

A cost model you can fill in

Use one scope, one period and one quality bar for every option. The diagram shows where the gain has to survive, and the table defines the inputs.

Where the gain has to surviveFour stages in a row: draft, review, test and fix, and ship. Drafting is where the agent speeds work up, which is the gain. Review, testing and later incidents are costs. The net gain is read at the ship stage. Below the stages, the break-even rule says the net gain must exceed the tool, review and quality costs divided by the loaded cost.Where the gain has to surviveDraftagent speeds this upReviewa human reads itTest and fixbugs surface laterShipmeasure it heregaincost Vcost Qnet gain gBreak-even: g must exceed (T + V + Q) / L
The gain is made at the draft and spent at review, testing and incidents. Measure it from ticket to production.
InputWhat to enterWhere it comes from
L, loaded costMonthly cost of the senior: salary, employer costs, equipment, overheadPayroll and finance
T, tool costSeats, usage above the plan, API spend, hosting for agentsVendor invoices
V, review costHours others spend reviewing and fixing agent output, times their hourly costPull request time logs
Q, quality costExpected monthly cost of incidents, rework and customer creditsIncident and defect log
B, baselineDays of scoped work delivered per month before agentsThree months of tracking
g, net gainMeasured gain from ticket to production at the same quality, as a fraction: 0.10 is 10%A pilot, not a survey
D and Sa, agencyAgency day rate, and the days it quotes for the same scopeWritten quote
cost per scope day, no agents      = L / B
cost per scope day, with agents    = (L + T + V + Q) / (B × (1 + g))
break-even net gain                = (T + V + Q) / L
senior with agents, scope S days   = (L + T + V + Q) × S / (B × (1 + g))
agency, same scope                 = D × Sa

The break-even line is the number to take into the room. Every point of tool, review and quality cost, as a share of the loaded cost, must be earned back as a point of measured net gain. If those costs add up to 10% of the loaded cost, a net gain below 10% makes each delivered day more expensive than before. For a team, add the loaded costs and use the team’s measured output.

The tool line is the easiest to estimate and the least stable. As of October 2026, Claude Pro costs $17 a month on an annual plan, or $20 billed monthly. A Claude Team standard seat costs $20 a month billed annually, a premium seat $100 a month billed annually, and Claude Max starts at $100 a month. GitHub Copilot Pro costs $10 a month, Pro+ $39 and Max $100. Usage limits apply, and Anthropic says its prices and plans may change at its discretion.

Compare the options on the same scope. The agency side is its day rate times its quoted days; the senior side is the scope formula. The answer flips in one of two ways: the measured gain is large and review cost is low, or the agency quotes far more days than the work needs. Ask both sides for the same deliverables before comparing numbers.

Three risks the spreadsheet will not show

Single point of failure

One senior is a bus factor of one, and agents deepen that dependency, because the know-how now sits in prompts, skills and configuration as well as in one head. Keep the agent configuration, the specs and the review rules in the repository. Name a second person who can review and ship, and agree in advance what happens during illness or holidays. The vendor is a second single point, so price a fallback, such as an agency retainer, and do not assume today’s price.

Quality debt

Quality debt arrives later and does not show up in coding time. DORA’s association with lower stability is the warning: code can arrive faster than the system absorbs it. The defence belongs in the pipeline, not the prompt. Make the merge depend on checks the agent cannot edit (my note on harness engineering for coding agents covers the set-up), require a human approval for every change, and track rework, such as changes reverted or reopened within a fixed window.

Data protection

Article 28 of the GDPR applies when a vendor processes personal data for you. The processor needs a written contract that sets out the subject-matter and duration of the processing, its nature and purpose, and the type of personal data. It may not bring in another processor without your prior written authorisation, specific or general. Put the agent vendor under that contract before personal data reaches a prompt, including names in tickets, customer records in test fixtures and personal data in logs.

Vendor terms matter too. Anthropic’s commercial terms say it may not train models on Customer Content from its services, and its Team plan lists no model training on your content by default. Customers in the EEA, Switzerland or the UK contract with Anthropic Ireland. Anthropic’s data-residency documentation describes US-only inference at 1.1 times the standard price, with global routing at standard pricing otherwise. In the page I read I found no EU-only option, so ask for one in writing. My note on GDPR and LLM API data residency covers the EU options in more detail.

When one senior with agents is enough

My rule of thumb is the chain below. Work through it in order and stop at the first no. The first three questions decide whether the setup is safe to run; the last one decides whether it is cheaper.

Deciding whether one senior with agents is enoughA chain of five steps, read from the top. If the scope is not written down, write the spec first. If tests do not catch regressions, fix the tests and CI first. If one person cannot review and ship, add a reviewer or a retainer. If the net gain is not above break-even, keep the team or hire the agency. If every check passes, take the work and re-measure every quarter.Start at the top, stop at the first noIs the scope written down?noWrite the spec firstA product owner or the team writes ityesDo tests catch regressions?noFix the tests and CI firstA failing test must stop the mergeyesCan a second person review and ship?noAdd a reviewer or a retainerNo single point of failureyesIs the net gain above break-even?noKeep the team or hire the agencyCompare the same scope and baryesTake it and re-measure quarterly
Stop at the first no. An unsafe setup is not cheaper, whatever the measured gain.

What I would do first

  1. Write down the baseline. For three months, record the days of scoped work this person delivers and the time from ticket to production.
  2. Run a two-week pilot on one bounded project with agents, and measure the end-to-end gain against the baseline, not the feeling of speed.
  3. Fill in the cost model with real numbers, and compare it with a written quote from an agency for the same scope.
  4. Sign the data processing agreement, and confirm training, retention and processing region in writing before client data goes into any prompt.
  5. Name a second person who reviews and ships agent work, and write down what happens when the senior is away.
  6. Re-measure every quarter. Tools change faster than the studies, and METR’s own follow-up shows how hard the number is to pin down.

None of this needs a platform. It needs a baseline, a measured gain and a signed contract, in that order.

Sources

  1. METR: early-2025 AI and experienced open-source developers
  2. Becker et al.: arXiv 2507.09089
  3. METR: developer productivity experiment design, February 2026
  4. Peng et al.: GitHub Copilot controlled experiment, arXiv 2302.06590
  5. Cui et al.: three field experiments with software developers
  6. DORA: State of AI-assisted Software Development 2025
  7. Google Cloud: highlights from the 2024 DORA report
  8. Stack Overflow: Developer Survey 2025, AI
  9. Ziftci et al.: Migrating Code At Scale With LLMs At Google
  10. Alshahwan et al.: Automated Unit Test Improvement using Large Language Models at Meta
  11. GitClear: AI Copilot Code Quality: 2025 Look Back at 12 Months of Data
  12. GDPR, Regulation (EU) 2016/679, Article 28
  13. Anthropic: Commercial Terms of Service
  14. Claude: pricing
  15. Claude Platform: data residency
  16. GitHub Copilot: plans and pricing

Frequently asked questions

Do AI coding agents make experienced developers faster?

The evidence does not settle it. A 2025 randomised trial by METR found that experienced open-source developers took 19% longer on tasks where AI was allowed, while they believed AI had made them 20% faster. METR’s February 2026 follow-up gave estimates whose confidence intervals include zero, and the authors called the data very weak evidence. Measure your own team before you decide.

What is the break-even productivity gain for an AI coding setup?

Add the monthly tool cost, the review time other people spend on the AI output and the expected monthly cost of quality problems. Divide that sum by the engineer’s fully loaded monthly cost. The result is the minimum net gain, measured end to end, that the setup must deliver just to break even.

Can I send client code or personal data to a coding agent under the GDPR?

Only under a written contract that meets Article 28, with the vendor as processor. Check in writing that the vendor does not train on your content, what it retains and where processing happens. Some business plans, such as Claude Team, include no model training by default, but the contract still has to be in place.

Is one senior with agents cheaper than an agency?

It can be, but the answer depends on the measured gain, the agency’s day rate and the days the agency needs for the same scope. Put both options on the same scope and quality bar, and compare the cost per delivered day. I found no public study that compares the two directly.

Sounds like what you need?

Tell me about your project or role – I’d love to hear from you.