Blog/AI agents

Multi-agent systems: when they beat one agent, and when they do not

Orchestrator-worker, fan-out, critic, handoff: what multi-agent systems really buy you, what they cost in tokens, how they fail, and a table to decide.

··12 min read

  • Multi-agent systems
  • AI agents
  • Orchestrator-worker
  • Context engineering
Diagram: a lead agent fans out to four worker agents, each with its own isolated context window, and gathers their summaries back.

Key takeaways

  • Anthropic measured about 4 times the tokens of a chat for an agent and about 15 times for a multi-agent system, so a multi-agent design has to be worth that multiplier.
  • The real benefit is context isolation: a worker explores in its own window and returns a summary, which keeps the lead agent focused and lets breadth exceed one context window.
  • Parallel reads scale well; parallel writes and tightly coupled steps do not. Google Research saw +81% on a parallelizable task and a 39 to 70% drop on a sequential one.
  • Most failures are coordination failures (vague specs, lost context, missing verification), not model failures, and debate-style setups often fail to beat a simple single-agent baseline.
  • Start with one agent and a good harness, add a subagent only for a measured reason, keep writes single-threaded, and evaluate the multi-agent version against the single-agent one.

Every few months a new framework makes it trivially easy to spin up a team of agents: a planner, a researcher, a coder, a reviewer. The diagrams look like an org chart, and the temptation is to assume that more agents means more intelligence. The published evidence says something more careful: multi-agent systems are excellent at a narrow class of problems, expensive everywhere, and quietly worse than a single agent on a surprising number of tasks.

This post sorts the patterns, puts numbers on the cost, lists the failure modes that the research and the vendors have documented, and ends with a decision framework you can apply to your own use case. My position, stated up front: start with one agent, and treat every additional agent as a purchase you must justify. The strongest justification is not "specialization" or "teamwork". It is context isolation.

If you want the single-agent basics first, read the agent loop explained. Everything below assumes you already have one agent that works.

Four patterns, and what each one is for

Most multi-agent designs are a combination of four shapes. They differ in who holds control, who sees which context, and where the results are merged.

Four multi-agent patternsFour small diagrams. Orchestrator-worker: a lead agent delegates to three workers. Parallel fan-out: one query goes to three searches whose results are merged. Writer and critic: a writer and a critic with a fresh context exchange drafts and feedback. Handoff: control passes from agent A to agent B to agent C.Four multi-agent patternsAnthropic, Cognition, OpenAIOrchestrator-workerLeadWorker 1Worker 2Worker 3Subtasks are not known in advanceParallel fan-outQuerySearch ASearch BSearch CMergeIndependent reads, one synthesisWriter and criticWriterCriticfresh contextIndependent check of an outputHandoffAgent AAgent BAgent CControl moves, one agent at a time
The four shapes. In the first two the caller stays in charge and receives summaries; in the handoff the control itself moves.

In more detail:

  • Orchestrator-worker. A lead agent plans, delegates to workers and synthesizes. Anthropic's Building effective agents describes it as a central LLM that "dynamically breaks down tasks, delegates them to worker LLMs, and synthesizes their results", suited to tasks where you cannot predict the subtasks in advance.
  • Parallel fan-out. The same idea with a fixed shape: split the work into independent parts, run them at once, merge. The same post names two variants, sectioning (independent subtasks) and voting (the same task run several times for diverse outputs).
  • Writer and critic (debate). A second agent checks the first one's output, or several agents argue towards an answer. This is where the evidence is most mixed, as the failure section shows.
  • Handoff. An agent passes the conversation to a specialist. In the OpenAI Agents SDK a handoff is exposed to the model as a tool named like transfer_to_refund_agent, and by default "the new agent takes over the conversation, and gets to see the entire previous conversation history". An input filter lets you narrow that.

Context isolation is the real benefit

Ask what a second agent can do that the first one cannot. It does not have a better model, and it does not think harder. What it has is an empty context window. A worker can read forty search results, grep a monorepo or run a noisy test suite, and hand back ten lines. The lead agent never carries that noise.

Anthropic's own description of its research system makes this explicit: subagents operate in parallel with their own context windows, which is how the system handles information that exceeds a single context window. Claude Code's documentation says the same for coding: use a subagent when a side task "would flood your main conversation with search results, logs, or file contents you won't reference again", because it does that work in its own context and "returns only the summary".

The flip side is just as important. A fresh subagent does not inherit your conversation history, your skills or the files already read, so everything it needs must be in its brief. The docs list when to stay in the main conversation: frequent back-and-forth, several phases that share significant context such as planning, implementation and testing, and latency-sensitive work. This is the same insight as in harness engineering for coding agents: what you put into the window, and what you keep out, decides the result.

Cognition found a second, subtler benefit: a clean context improves a reviewer. In their April 2026 follow-up they report that their review agent works better when it does not share context with the coding agent, because it reasons independently instead of inheriting the author's assumptions. They state that Devin Review catches an average of 2 bugs per pull request, about 58% of them severe. That is a vendor figure, but the mechanism is plausible and cheap to test. It also fits the review bottleneck described in AI-generated pull requests.

What it costs: the token multiplier

Anthropic is unusually candid about this in its multi-agent research system post: agents typically use about 4 times more tokens than chat interactions, and multi-agent systems about 15 times more. They conclude that such systems need tasks whose value justifies the cost.

There is an uncomfortable reading of the same post. In their analysis of the BrowseComp benchmark, token usage by itself explained 80% of the variance in performance. Part of what a multi-agent system buys is simply more thinking per question. That is legitimate, but it means you should compare against a single agent given the same token budget, not against a single agent that stops early. Their headline result is that a Claude Opus 4 lead with Claude Sonnet 4 subagents beat a single Claude Opus 4 by 90.2% on their internal research evaluation, with parallelization cutting research time by up to 90% for complex queries.

Latency and money pull in opposite directions. Parallel workers shorten wall-clock time and raise the bill. Per-token price falls with caching and routing, so read LLM cost, latency, prompt caching and routing before you conclude that the multiplier is unaffordable. Cheaper worker models, shared cached prefixes and strict caps on the number of workers change the economics a lot.

Anthropic also lists the failure modes of its early versions: spawning 50 subagents for a simple query, searching endlessly for sources that do not exist, and workers duplicating each other's work because the task descriptions were vague. Their fix was explicit scaling rules in the prompt, so that effort matches the complexity of the query, and much more detailed task descriptions for every worker.

How multi-agent systems fail

The research is consistent on one point: the failures are mostly about coordination, not about model intelligence.

  • Dispersed decisions. Cognition's June 2025 post "Don't Build Multi-Agents" rests on two principles: share context, and "actions carry implicit decisions". Their example is a Flappy Bird clone split into subtasks, where one subagent builds a Super Mario style background and another a bird that does not look or behave like Flappy Bird, and the final agent has to merge the mismatch. Their summary: running multiple agents in collaboration only results in fragile systems.
  • Sequential work gets worse. Google Research evaluated 180 agent configurations. Centralized coordination improved a parallelizable financial reasoning task by 80.9% over a single agent, while on a sequential planning task every multi-agent variant tested degraded performance by 39 to 70%. Independent agents amplified errors 17.2 times, a central orchestrator only 4.4 times, because it acts as a validation bottleneck. The authors also report a tool-coordination trade-off: overhead grows disproportionately for tool-heavy tasks.
  • Taxonomy of breakdowns. The MAST study (Cemri et al.) analysed more than 1,600 annotated traces from 7 multi-agent frameworks and grouped 14 failure modes into three categories: system design issues, inter-agent misalignment and task verification. The authors note that performance gains on popular benchmarks are often minimal, and in several cases the same model in a single-agent setup did better.
  • Debate is overrated by default. Du et al. (2023) showed that multiple model instances debating can improve reasoning and factuality. A 2025 evaluation of 5 debate methods across 9 benchmarks and 4 models then found that they often fail to outperform Chain-of-Thought and Self-Consistency, even with much more inference-time compute. What helped was model heterogeneity: debaters from different models.
  • Parallel writers conflict. In April 2026 Cognition updated its view: multi-agent systems work best today when writes stay single-threaded and the additional agents contribute intelligence rather than actions, and most swarm-style ideas still see little adoption. Anthropic likewise calls most coding work a poor fit because of limited parallelization.

The pattern across these sources is a split between reading and writing. Reading, searching, analysing and reviewing parallelize well, because each result can be judged on its own. Writing code, editing shared state and making design choices do not, because every action carries decisions the other agents cannot see.

Verification is the other recurring gap. If nobody checks the merged result, errors propagate; Google's numbers show how much a central validation step contains. Treat your orchestrator's merge step as a place for explicit checks, and measure the whole system with evals built for the feature, not by reading a few traces.

A decision framework

Here is the table I would use in a design review. The question is never "single or multi", but which specific shape pays for itself on this task.

SituationSignalPatternWhy
Broad research over many independent sourcesSubtasks are unknown upfront and exceed one context windowOrchestrator-workerIsolated windows give breadth; Anthropic reports large gains here at roughly 15 times the tokens of a chat
Known set of independent checks or lookupsSame operation over N items, no dependenciesParallel fan-out (sectioning)Wall-clock time drops, no coordination needed beyond a merge
High-stakes output that a second look could catchReview benefits from not sharing the author's assumptionsWriter and critic with a fresh contextIndependent review; ideally a different model, as the debate study suggests
Several domains with different tools or promptsRouting between specialists, one active at a timeHandoffEach specialist gets a short prompt and few tools; decide how much history moves
A side task with noisy outputLogs, search results or file contents you will not reference againSubagent that returns a summaryKeeps the main context clean at the price of a fresh start
Multi-step work where steps depend on each otherPlanning, refactoring, one document, shared stateSingle agentSequential tasks degraded by 39 to 70% in multi-agent variants in Google's study
Parallel edits to the same codebaseConflicting style and edge-case decisionsSingle writer, helpers read onlyCognition: keep writes single-threaded

If a row says single agent, the usual fix for a struggling system is not another agent but a better harness: clearer instructions, better tools, compaction and checkpoints. Anthropic's own advice in Building effective agents is to find the simplest solution possible and increase complexity only when it demonstrably improves outcomes, and it warns that frameworks can obscure prompts and responses and make debugging harder.

A checklist before you add an agent

  1. Build the single-agent baseline first, with the same tools and a fair token budget.
  2. Write down the isolation argument: what stays out of whose context.
  3. Classify the work as read-heavy or write-heavy. Keep writes single-threaded.
  4. Give every worker a full brief: objective, output format, tools, boundaries and a stop condition, since it inherits nothing.
  5. Put explicit scaling rules in the lead prompt and a hard cap on workers and turns.
  6. Add a verification step on the merged result, and decide who is allowed to say "done".
  7. Evaluate both versions on about 20 realistic queries first, as Anthropic suggests, then grow the set. Compare quality, tokens, latency and failure rate.
  8. Log every delegation with its brief and its summary so you can debug lost context, and plan how running agents survive a deployment.

On the last point, Anthropic notes that agent processes are stateful and long-running, so it uses rainbow deployments that shift traffic gradually instead of interrupting runs in progress. Cost control is part of the same discipline; the tools in token-saving tools for coding agents apply to workers too.

What I would do

For a typical company use case, I would ship a single agent with strong tools and a good harness, and add exactly one multi-agent element where the isolation argument is strong: a research subagent that returns cited summaries, or a clean-context reviewer on the output. Both keep the writes with one agent, and both are easy to measure against the baseline.

I would avoid swarms of peers, open-ended debate between same-model agents, and parallel code writers until the evidence changes. Even Cognition, the loudest critic, now describes a narrower class of multi-agent designs that work. The recent vendor posts I read describe the same shape: one orchestrator that owns the context, with isolated helpers that return summaries. That is the architecture worth learning, and it is a good fit for the kind of AI engineering work I do for clients.

Sources

  1. Anthropic Engineering: How we built our multi-agent research system
  2. Anthropic: Building effective agents
  3. Cognition: Don't Build Multi-Agents (12 June 2025)
  4. Cognition: Multi-Agents: What's Actually Working (22 April 2026)
  5. Google Research: Towards a science of scaling agent systems
  6. arXiv 2512.08296: Towards a Science of Scaling Agent Systems
  7. arXiv 2503.13657: Why Do Multi-Agent LLM Systems Fail? (MAST)
  8. arXiv 2305.14325: Improving Factuality and Reasoning in Language Models through Multiagent Debate
  9. arXiv 2502.08788: Stop Overvaluing Multi-Agent Debate
  10. OpenAI Agents SDK: Handoffs
  11. Claude Code documentation: Subagents

Frequently asked questions

When should I use a multi-agent system instead of a single agent?

When the work is wide rather than deep: many independent questions, sources that exceed one context window, or side tasks that would flood the main conversation with output. Anthropic describes breadth-first research as the sweet spot. If the steps depend on each other or share a lot of context, a single agent is usually better.

How many more tokens do multi-agent systems use?

Anthropic reported that agents use about 4 times more tokens than chat interactions and multi-agent systems about 15 times more. Their analysis of the BrowseComp benchmark found that token usage alone explained about 80% of the performance variance, so part of the gain is simply more compute.

Why do multi-agent LLM systems fail?

The MAST study of more than 1,600 annotated traces from 7 frameworks groups 14 failure modes into system design issues, inter-agent misalignment and task verification. In practice that means vague task descriptions, context that does not reach the next agent, duplicated work and nobody checking the final result.

Is multi-agent debate worth it?

Rarely as a default. The original debate paper reported better reasoning and factuality, but a later evaluation of 5 debate methods on 9 benchmarks found they often fail to beat Chain-of-Thought or self-consistency despite using more compute. Using different models as debaters helped in that study.

What is the difference between a handoff and a subagent?

A handoff transfers control: the next agent takes over the conversation, by default with the full history. A subagent is delegated a side task in a fresh context and returns only a summary while the caller stays in charge. Handoffs suit routing between specialists, subagents suit context isolation.

Should coding agents run subagents in parallel?

Be careful. Cognition argues that parallel writers make conflicting implicit decisions about style and edge cases, and says multi-agent setups work best today when writes stay single-threaded. Subagents for exploration, search and review are the safe use.

Sounds like what you need?

Tell me about your project or role – I’d love to hear from you.