[{"data":1,"prerenderedAt":754},["ShallowReactive",2],{"blog-multi-agent-systems-when-worth-it-en":3},{"slug":4,"published":5,"minutes":6,"category":7,"tags":8,"keywords":13,"about":24,"sources":34,"cover":68,"og":69,"expertise":70,"locales":71,"lang":72,"title":75,"description":76,"coverAlt":77,"metaTitle":78,"takeaways":79,"faq":85,"toc":104,"blocks":129,"others":456},"multi-agent-systems-when-worth-it","2026-10-02",12,"agents",[9,10,11,12],"Multi-agent systems","AI agents","Orchestrator-worker","Context engineering",[14,15,16,17,18,19,20,21,22,23],"multi-agent systems when worth it","multi-agent vs single agent","orchestrator-worker pattern","multi-agent LLM token cost","why multi-agent systems fail","don't build multi-agents","subagents context isolation","multi-agent debate worth it","agent handoffs vs subagents","when to use multiple AI agents",[25,28,31],{"name":26,"url":27},"Multi-agent system","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FMulti-agent_system",{"name":29,"url":30},"Intelligent agent","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FIntelligent_agent",{"name":32,"url":33},"Large language model","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FLarge_language_model",[35,38,41,44,47,50,53,56,59,62,65],{"title":36,"url":37},"Anthropic Engineering: How we built our multi-agent research system","https:\u002F\u002Fwww.anthropic.com\u002Fengineering\u002Fmulti-agent-research-system",{"title":39,"url":40},"Anthropic: Building effective agents","https:\u002F\u002Fwww.anthropic.com\u002Fengineering\u002Fbuilding-effective-agents",{"title":42,"url":43},"Cognition: Don't Build Multi-Agents (12 June 2025)","https:\u002F\u002Fcognition.com\u002Fblog\u002Fdont-build-multi-agents",{"title":45,"url":46},"Cognition: Multi-Agents: What's Actually Working (22 April 2026)","https:\u002F\u002Fcognition.com\u002Fblog\u002Fmulti-agents-working",{"title":48,"url":49},"Google Research: Towards a science of scaling agent systems","https:\u002F\u002Fresearch.google\u002Fblog\u002Ftowards-a-science-of-scaling-agent-systems-when-and-why-agent-systems-work\u002F",{"title":51,"url":52},"arXiv 2512.08296: Towards a Science of Scaling Agent Systems","https:\u002F\u002Farxiv.org\u002Fabs\u002F2512.08296",{"title":54,"url":55},"arXiv 2503.13657: Why Do Multi-Agent LLM Systems Fail? (MAST)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2503.13657",{"title":57,"url":58},"arXiv 2305.14325: Improving Factuality and Reasoning in Language Models through Multiagent Debate","https:\u002F\u002Farxiv.org\u002Fabs\u002F2305.14325",{"title":60,"url":61},"arXiv 2502.08788: Stop Overvaluing Multi-Agent Debate","https:\u002F\u002Farxiv.org\u002Fabs\u002F2502.08788",{"title":63,"url":64},"OpenAI Agents SDK: Handoffs","https:\u002F\u002Fopenai.github.io\u002Fopenai-agents-python\u002Fhandoffs\u002F",{"title":66,"url":67},"Claude Code documentation: Subagents","https:\u002F\u002Fcode.claude.com\u002Fdocs\u002Fen\u002Fsub-agents","\u002Fimages\u002Fblog\u002Fmulti-agent-systems-when-worth-it\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fmulti-agent-systems-when-worth-it\u002Fog.jpg","ai-engineer",[72,73,74],"en","de","hu","Multi-agent systems: when they beat one agent, and when they do not","Orchestrator-worker, fan-out, critic, handoff: what multi-agent systems really buy you, what they cost in tokens, how they fail, and a table to decide.","Diagram: a lead agent fans out to four worker agents, each with its own isolated context window, and gathers their summaries back.","When multi-agent systems beat one agent · Balázs Csorba",[80,81,82,83,84],"Anthropic measured about 4 times the tokens of a chat for an agent and about 15 times for a multi-agent system, so a multi-agent design has to be worth that multiplier.","The real benefit is context isolation: a worker explores in its own window and returns a summary, which keeps the lead agent focused and lets breadth exceed one context window.","Parallel reads scale well; parallel writes and tightly coupled steps do not. Google Research saw +81% on a parallelizable task and a 39 to 70% drop on a sequential one.","Most failures are coordination failures (vague specs, lost context, missing verification), not model failures, and debate-style setups often fail to beat a simple single-agent baseline.","Start with one agent and a good harness, add a subagent only for a measured reason, keep writes single-threaded, and evaluate the multi-agent version against the single-agent one.",[86,89,92,95,98,101],{"q":87,"a":88},"When should I use a multi-agent system instead of a single agent?","When the work is wide rather than deep: many independent questions, sources that exceed one context window, or side tasks that would flood the main conversation with output. Anthropic describes breadth-first research as the sweet spot. If the steps depend on each other or share a lot of context, a single agent is usually better.",{"q":90,"a":91},"How many more tokens do multi-agent systems use?","Anthropic reported that agents use about 4 times more tokens than chat interactions and multi-agent systems about 15 times more. Their analysis of the BrowseComp benchmark found that token usage alone explained about 80% of the performance variance, so part of the gain is simply more compute.",{"q":93,"a":94},"Why do multi-agent LLM systems fail?","The MAST study of more than 1,600 annotated traces from 7 frameworks groups 14 failure modes into system design issues, inter-agent misalignment and task verification. In practice that means vague task descriptions, context that does not reach the next agent, duplicated work and nobody checking the final result.",{"q":96,"a":97},"Is multi-agent debate worth it?","Rarely as a default. The original debate paper reported better reasoning and factuality, but a later evaluation of 5 debate methods on 9 benchmarks found they often fail to beat Chain-of-Thought or self-consistency despite using more compute. Using different models as debaters helped in that study.",{"q":99,"a":100},"What is the difference between a handoff and a subagent?","A handoff transfers control: the next agent takes over the conversation, by default with the full history. A subagent is delegated a side task in a fresh context and returns only a summary while the caller stays in charge. Handoffs suit routing between specialists, subagents suit context isolation.",{"q":102,"a":103},"Should coding agents run subagents in parallel?","Be careful. Cognition argues that parallel writers make conflicting implicit decisions about style and edge cases, and says multi-agent setups work best today when writes stay single-threaded. Subagents for exploration, search and review are the safe use.",[105,108,111,114,117,120,123,126],{"id":106,"title":107},"four-patterns","Four patterns, and what each one is for",{"id":109,"title":110},"context-isolation","Context isolation is the real benefit",{"id":112,"title":113},"token-cost","What it costs: the token multiplier",{"id":115,"title":116},"coordination-failures","How multi-agent systems fail",{"id":118,"title":119},"decision-framework","A decision framework",{"id":121,"title":122},"checklist","A checklist before you add an agent",{"id":124,"title":125},"what-i-would-do","What I would do",{"id":127,"title":128},"sources","Sources",[130,134,142,151,154,157,166,169,198,199,206,209,217,224,232,233,236,243,251,254,255,258,285,292,300,301,304,378,381,382,401,409,410,413,420,421],{"type":131,"content":132},"paragraph",[133],"Every few months a new framework makes it trivially easy to spin up a team of agents: a planner, a researcher, a coder, a reviewer. The diagrams look like an org chart, and the temptation is to assume that more agents means more intelligence. The published evidence says something more careful: multi-agent systems are excellent at a narrow class of problems, expensive everywhere, and quietly worse than a single agent on a surprising number of tasks.",{"type":131,"content":135},[136,137,141],"This post sorts the patterns, puts numbers on the cost, lists the failure modes that the research and the vendors have documented, and ends with a decision framework you can apply to your own use case. My position, stated up front: ",{"tag":138,"children":139},"strong",[140],"start with one agent, and treat every additional agent as a purchase you must justify",". The strongest justification is not \"specialization\" or \"teamwork\". It is context isolation.",{"type":131,"content":143},[144,145,150],"If you want the single-agent basics first, read ",{"tag":146,"to":147,"children":148},"link","\u002Fblog\u002Fagent-loop-explained",[149],"the agent loop explained",". Everything below assumes you already have one agent that works.",{"type":152,"level":153,"id":106,"text":107},"heading",2,{"type":131,"content":155},[156],"Most multi-agent designs are a combination of four shapes. They differ in who holds control, who sees which context, and where the results are merged.",{"type":158,"attrs":159,"inner":163,"caption":164},"diagram",{"viewBox":160,"role":161,"aria-labelledby":162},"0 0 720 370","img","d1-multi-t d1-multi-d","\u003Ctitle id=\"d1-multi-t\">Four multi-agent patterns\u003C\u002Ftitle>\u003Cdesc id=\"d1-multi-d\">Four small diagrams. Orchestrator-worker: a lead agent delegates to three workers. Parallel fan-out: one query goes to three searches whose results are merged. Writer and critic: a writer and a critic with a fresh context exchange drafts and feedback. Handoff: control passes from agent A to agent B to agent C.\u003C\u002Fdesc>\u003Ctext x=\"20\" y=\"26\" class=\"d-title\">Four multi-agent patterns\u003C\u002Ftext>\u003Ctext x=\"700\" y=\"26\" text-anchor=\"end\" class=\"d-label\">Anthropic, Cognition, OpenAI\u003C\u002Ftext>\u003Crect x=\"20\" y=\"40\" width=\"340\" height=\"150\" rx=\"12\" class=\"d-box d-dash\" \u002F>\u003Ctext x=\"34\" y=\"62\" class=\"d-label\">Orchestrator-worker\u003C\u002Ftext>\u003Crect x=\"34\" y=\"98\" width=\"90\" height=\"36\" rx=\"8\" class=\"d-accent\" \u002F>\u003Ctext x=\"79\" y=\"120\" text-anchor=\"middle\" class=\"d-small\">Lead\u003C\u002Ftext>\u003Cpath d=\"M124 116 L213 89\" class=\"d-line\" \u002F>\u003Cpath d=\"M220 89 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Crect x=\"220\" y=\"76\" width=\"126\" height=\"26\" rx=\"8\" class=\"d-sky\" \u002F>\u003Ctext x=\"283\" y=\"93\" text-anchor=\"middle\" class=\"d-small\">Worker 1\u003C\u002Ftext>\u003Cpath d=\"M124 116 L213 121\" class=\"d-line\" \u002F>\u003Cpath d=\"M220 121 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Crect x=\"220\" y=\"108\" width=\"126\" height=\"26\" rx=\"8\" class=\"d-sky\" \u002F>\u003Ctext x=\"283\" y=\"125\" text-anchor=\"middle\" class=\"d-small\">Worker 2\u003C\u002Ftext>\u003Cpath d=\"M124 116 L213 153\" class=\"d-line\" \u002F>\u003Cpath d=\"M220 153 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Crect x=\"220\" y=\"140\" width=\"126\" height=\"26\" rx=\"8\" class=\"d-sky\" \u002F>\u003Ctext x=\"283\" y=\"157\" text-anchor=\"middle\" class=\"d-small\">Worker 3\u003C\u002Ftext>\u003Ctext x=\"190\" y=\"180\" text-anchor=\"middle\" class=\"d-small\">Subtasks are not known in advance\u003C\u002Ftext>\u003Crect x=\"360\" y=\"40\" width=\"340\" height=\"150\" rx=\"12\" class=\"d-box d-dash\" \u002F>\u003Ctext x=\"374\" y=\"62\" class=\"d-label\">Parallel fan-out\u003C\u002Ftext>\u003Crect x=\"372\" y=\"98\" width=\"70\" height=\"36\" rx=\"8\" class=\"d-box\" \u002F>\u003Ctext x=\"407\" y=\"120\" text-anchor=\"middle\" class=\"d-small\">Query\u003C\u002Ftext>\u003Cpath d=\"M442 116 L465 89\" class=\"d-line\" \u002F>\u003Cpath d=\"M472 89 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Crect x=\"472\" y=\"76\" width=\"100\" height=\"26\" rx=\"8\" class=\"d-sky\" \u002F>\u003Ctext x=\"522\" y=\"93\" text-anchor=\"middle\" class=\"d-small\">Search A\u003C\u002Ftext>\u003Cpath d=\"M572 89 L593 116\" class=\"d-line\" \u002F>\u003Cpath d=\"M600 116 l-8 -3 l2 7 z\" class=\"d-head\" \u002F>\u003Cpath d=\"M442 116 L465 121\" class=\"d-line\" \u002F>\u003Cpath d=\"M472 121 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Crect x=\"472\" y=\"108\" width=\"100\" height=\"26\" rx=\"8\" class=\"d-sky\" \u002F>\u003Ctext x=\"522\" y=\"125\" text-anchor=\"middle\" class=\"d-small\">Search B\u003C\u002Ftext>\u003Cpath d=\"M572 121 L593 116\" class=\"d-line\" \u002F>\u003Cpath d=\"M600 116 l-8 -3 l2 7 z\" class=\"d-head\" \u002F>\u003Cpath d=\"M442 116 L465 153\" class=\"d-line\" \u002F>\u003Cpath d=\"M472 153 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Crect x=\"472\" y=\"140\" width=\"100\" height=\"26\" rx=\"8\" class=\"d-sky\" \u002F>\u003Ctext x=\"522\" y=\"157\" text-anchor=\"middle\" class=\"d-small\">Search C\u003C\u002Ftext>\u003Cpath d=\"M572 153 L593 116\" class=\"d-line\" \u002F>\u003Cpath d=\"M600 116 l-8 -3 l2 7 z\" class=\"d-head\" \u002F>\u003Crect x=\"600\" y=\"98\" width=\"88\" height=\"36\" rx=\"8\" class=\"d-mint\" \u002F>\u003Ctext x=\"644\" y=\"120\" text-anchor=\"middle\" class=\"d-small\">Merge\u003C\u002Ftext>\u003Ctext x=\"530\" y=\"180\" text-anchor=\"middle\" class=\"d-small\">Independent reads, one synthesis\u003C\u002Ftext>\u003Crect x=\"20\" y=\"205\" width=\"340\" height=\"150\" rx=\"12\" class=\"d-box d-dash\" \u002F>\u003Ctext x=\"34\" y=\"227\" class=\"d-label\">Writer and critic\u003C\u002Ftext>\u003Crect x=\"40\" y=\"257\" width=\"110\" height=\"44\" rx=\"8\" class=\"d-box\" \u002F>\u003Ctext x=\"95\" y=\"283\" text-anchor=\"middle\" class=\"d-small\">Writer\u003C\u002Ftext>\u003Crect x=\"230\" y=\"257\" width=\"110\" height=\"44\" rx=\"8\" class=\"d-gold\" \u002F>\u003Ctext x=\"285\" y=\"283\" text-anchor=\"middle\" class=\"d-small\">Critic\u003C\u002Ftext>\u003Cpath d=\"M150 269 H222\" class=\"d-line\" \u002F>\u003Cpath d=\"M230 269 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Cpath d=\"M230 291 H158\" class=\"d-line\" \u002F>\u003Cpath d=\"M150 291 l9 -5 v10 z\" class=\"d-head\" \u002F>\u003Ctext x=\"285\" y=\"321\" text-anchor=\"middle\" class=\"d-label\">fresh context\u003C\u002Ftext>\u003Ctext x=\"190\" y=\"345\" text-anchor=\"middle\" class=\"d-small\">Independent check of an output\u003C\u002Ftext>\u003Crect x=\"360\" y=\"205\" width=\"340\" height=\"150\" rx=\"12\" class=\"d-box d-dash\" \u002F>\u003Ctext x=\"374\" y=\"227\" class=\"d-label\">Handoff\u003C\u002Ftext>\u003Crect x=\"372\" y=\"263\" width=\"86\" height=\"40\" rx=\"8\" class=\"d-sky\" \u002F>\u003Ctext x=\"415\" y=\"287\" text-anchor=\"middle\" class=\"d-small\">Agent A\u003C\u002Ftext>\u003Cpath d=\"M458 283 H478\" class=\"d-line\" \u002F>\u003Cpath d=\"M486 283 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Crect x=\"486\" y=\"263\" width=\"86\" height=\"40\" rx=\"8\" class=\"d-accent\" \u002F>\u003Ctext x=\"529\" y=\"287\" text-anchor=\"middle\" class=\"d-small\">Agent B\u003C\u002Ftext>\u003Cpath d=\"M572 283 H592\" class=\"d-line\" \u002F>\u003Cpath d=\"M600 283 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Crect x=\"600\" y=\"263\" width=\"88\" height=\"40\" rx=\"8\" class=\"d-mint\" \u002F>\u003Ctext x=\"644\" y=\"287\" text-anchor=\"middle\" class=\"d-small\">Agent C\u003C\u002Ftext>\u003Ctext x=\"530\" y=\"345\" text-anchor=\"middle\" class=\"d-small\">Control moves, one agent at a time\u003C\u002Ftext>",[165],"The four shapes. In the first two the caller stays in charge and receives summaries; in the handoff the control itself moves.",{"type":131,"content":167},[168],"In more detail:",{"type":170,"ordered":171,"items":172},"list",false,[173,183,188,193],[174,177,178,182],{"tag":138,"children":175},[176],"Orchestrator-worker."," A lead agent plans, delegates to workers and synthesizes. Anthropic's ",{"tag":179,"href":40,"children":180},"a",[181],"Building effective agents"," describes it as a central LLM that \"dynamically breaks down tasks, delegates them to worker LLMs, and synthesizes their results\", suited to tasks where you cannot predict the subtasks in advance.",[184,187],{"tag":138,"children":185},[186],"Parallel fan-out."," The same idea with a fixed shape: split the work into independent parts, run them at once, merge. The same post names two variants, sectioning (independent subtasks) and voting (the same task run several times for diverse outputs).",[189,192],{"tag":138,"children":190},[191],"Writer and critic (debate)."," A second agent checks the first one's output, or several agents argue towards an answer. This is where the evidence is most mixed, as the failure section shows.",[194,197],{"tag":138,"children":195},[196],"Handoff."," An agent passes the conversation to a specialist. In the OpenAI Agents SDK a handoff is exposed to the model as a tool named like transfer_to_refund_agent, and by default \"the new agent takes over the conversation, and gets to see the entire previous conversation history\". An input filter lets you narrow that.",{"type":152,"level":153,"id":109,"text":110},{"type":131,"content":200},[201,202,205],"Ask what a second agent can do that the first one cannot. It does not have a better model, and it does not think harder. What it has is ",{"tag":138,"children":203},[204],"an empty context window",". A worker can read forty search results, grep a monorepo or run a noisy test suite, and hand back ten lines. The lead agent never carries that noise.",{"type":131,"content":207},[208],"Anthropic's own description of its research system makes this explicit: subagents operate in parallel with their own context windows, which is how the system handles information that exceeds a single context window. Claude Code's documentation says the same for coding: use a subagent when a side task \"would flood your main conversation with search results, logs, or file contents you won't reference again\", because it does that work in its own context and \"returns only the summary\".",{"type":131,"content":210},[211,212,216],"The flip side is just as important. A fresh subagent does not inherit your conversation history, your skills or the files already read, so everything it needs must be in its brief. The docs list when to stay in the main conversation: frequent back-and-forth, several phases that share significant context such as planning, implementation and testing, and latency-sensitive work. This is the same insight as in ",{"tag":146,"to":213,"children":214},"\u002Fblog\u002Fharness-engineering-coding-agents",[215],"harness engineering for coding agents",": what you put into the window, and what you keep out, decides the result.",{"type":218,"variant":219,"title":220,"body":221},"callout","tip","A useful test",[222],[223],"Before adding an agent, finish this sentence: \"This agent exists so that ___ never enters the other agent's context.\" If you cannot fill the blank, you are probably adding coordination cost without isolation benefit.",{"type":131,"content":225},[226,227,231],"Cognition found a second, subtler benefit: a clean context improves a reviewer. In their April 2026 follow-up they report that their review agent works better when it does not share context with the coding agent, because it reasons independently instead of inheriting the author's assumptions. They state that Devin Review catches an average of 2 bugs per pull request, about 58% of them severe. That is a vendor figure, but the mechanism is plausible and cheap to test. It also fits the review bottleneck described in ",{"tag":146,"to":228,"children":229},"\u002Fblog\u002Fai-generated-pr-review-bottleneck",[230],"AI-generated pull requests",".",{"type":152,"level":153,"id":112,"text":113},{"type":131,"content":234},[235],"Anthropic is unusually candid about this in its multi-agent research system post: agents typically use about 4 times more tokens than chat interactions, and multi-agent systems about 15 times more. They conclude that such systems need tasks whose value justifies the cost.",{"type":131,"content":237},[238,239,242],"There is an uncomfortable reading of the same post. In their analysis of the BrowseComp benchmark, token usage by itself explained 80% of the variance in performance. Part of what a multi-agent system buys is simply more thinking per question. That is legitimate, but it means you should compare against a ",{"tag":138,"children":240},[241],"single agent given the same token budget",", not against a single agent that stops early. Their headline result is that a Claude Opus 4 lead with Claude Sonnet 4 subagents beat a single Claude Opus 4 by 90.2% on their internal research evaluation, with parallelization cutting research time by up to 90% for complex queries.",{"type":131,"content":244},[245,246,250],"Latency and money pull in opposite directions. Parallel workers shorten wall-clock time and raise the bill. Per-token price falls with caching and routing, so read ",{"tag":146,"to":247,"children":248},"\u002Fblog\u002Fllm-cost-latency-prompt-caching-routing",[249],"LLM cost, latency, prompt caching and routing"," before you conclude that the multiplier is unaffordable. Cheaper worker models, shared cached prefixes and strict caps on the number of workers change the economics a lot.",{"type":131,"content":252},[253],"Anthropic also lists the failure modes of its early versions: spawning 50 subagents for a simple query, searching endlessly for sources that do not exist, and workers duplicating each other's work because the task descriptions were vague. Their fix was explicit scaling rules in the prompt, so that effort matches the complexity of the query, and much more detailed task descriptions for every worker.",{"type":152,"level":153,"id":115,"text":116},{"type":131,"content":256},[257],"The research is consistent on one point: the failures are mostly about coordination, not about model intelligence.",{"type":170,"ordered":171,"items":259},[260,265,270,275,280],[261,264],{"tag":138,"children":262},[263],"Dispersed decisions."," Cognition's June 2025 post \"Don't Build Multi-Agents\" rests on two principles: share context, and \"actions carry implicit decisions\". Their example is a Flappy Bird clone split into subtasks, where one subagent builds a Super Mario style background and another a bird that does not look or behave like Flappy Bird, and the final agent has to merge the mismatch. Their summary: running multiple agents in collaboration only results in fragile systems.",[266,269],{"tag":138,"children":267},[268],"Sequential work gets worse."," Google Research evaluated 180 agent configurations. Centralized coordination improved a parallelizable financial reasoning task by 80.9% over a single agent, while on a sequential planning task every multi-agent variant tested degraded performance by 39 to 70%. Independent agents amplified errors 17.2 times, a central orchestrator only 4.4 times, because it acts as a validation bottleneck. The authors also report a tool-coordination trade-off: overhead grows disproportionately for tool-heavy tasks.",[271,274],{"tag":138,"children":272},[273],"Taxonomy of breakdowns."," The MAST study (Cemri et al.) analysed more than 1,600 annotated traces from 7 multi-agent frameworks and grouped 14 failure modes into three categories: system design issues, inter-agent misalignment and task verification. The authors note that performance gains on popular benchmarks are often minimal, and in several cases the same model in a single-agent setup did better.",[276,279],{"tag":138,"children":277},[278],"Debate is overrated by default."," Du et al. (2023) showed that multiple model instances debating can improve reasoning and factuality. A 2025 evaluation of 5 debate methods across 9 benchmarks and 4 models then found that they often fail to outperform Chain-of-Thought and Self-Consistency, even with much more inference-time compute. What helped was model heterogeneity: debaters from different models.",[281,284],{"tag":138,"children":282},[283],"Parallel writers conflict."," In April 2026 Cognition updated its view: multi-agent systems work best today when writes stay single-threaded and the additional agents contribute intelligence rather than actions, and most swarm-style ideas still see little adoption. Anthropic likewise calls most coding work a poor fit because of limited parallelization.",{"type":131,"content":286},[287,288,291],"The pattern across these sources is a split between ",{"tag":138,"children":289},[290],"reading and writing",". Reading, searching, analysing and reviewing parallelize well, because each result can be judged on its own. Writing code, editing shared state and making design choices do not, because every action carries decisions the other agents cannot see.",{"type":131,"content":293},[294,295,299],"Verification is the other recurring gap. If nobody checks the merged result, errors propagate; Google's numbers show how much a central validation step contains. Treat your orchestrator's merge step as a place for explicit checks, and measure the whole system with ",{"tag":146,"to":296,"children":297},"\u002Fblog\u002Fllm-evals-for-product-features",[298],"evals built for the feature",", not by reading a few traces.",{"type":152,"level":153,"id":118,"text":119},{"type":131,"content":302},[303],"Here is the table I would use in a design review. The question is never \"single or multi\", but which specific shape pays for itself on this task.",{"type":305,"head":306,"rows":315},"table",[307,309,311,313],[308],"Situation",[310],"Signal",[312],"Pattern",[314],"Why",[316,324,333,342,351,360,369],[317,319,321,322],[318],"Broad research over many independent sources",[320],"Subtasks are unknown upfront and exceed one context window",[11],[323],"Isolated windows give breadth; Anthropic reports large gains here at roughly 15 times the tokens of a chat",[325,327,329,331],[326],"Known set of independent checks or lookups",[328],"Same operation over N items, no dependencies",[330],"Parallel fan-out (sectioning)",[332],"Wall-clock time drops, no coordination needed beyond a merge",[334,336,338,340],[335],"High-stakes output that a second look could catch",[337],"Review benefits from not sharing the author's assumptions",[339],"Writer and critic with a fresh context",[341],"Independent review; ideally a different model, as the debate study suggests",[343,345,347,349],[344],"Several domains with different tools or prompts",[346],"Routing between specialists, one active at a time",[348],"Handoff",[350],"Each specialist gets a short prompt and few tools; decide how much history moves",[352,354,356,358],[353],"A side task with noisy output",[355],"Logs, search results or file contents you will not reference again",[357],"Subagent that returns a summary",[359],"Keeps the main context clean at the price of a fresh start",[361,363,365,367],[362],"Multi-step work where steps depend on each other",[364],"Planning, refactoring, one document, shared state",[366],"Single agent",[368],"Sequential tasks degraded by 39 to 70% in multi-agent variants in Google's study",[370,372,374,376],[371],"Parallel edits to the same codebase",[373],"Conflicting style and edge-case decisions",[375],"Single writer, helpers read only",[377],"Cognition: keep writes single-threaded",{"type":131,"content":379},[380],"If a row says single agent, the usual fix for a struggling system is not another agent but a better harness: clearer instructions, better tools, compaction and checkpoints. Anthropic's own advice in Building effective agents is to find the simplest solution possible and increase complexity only when it demonstrably improves outcomes, and it warns that frameworks can obscure prompts and responses and make debugging harder.",{"type":152,"level":153,"id":121,"text":122},{"type":170,"ordered":383,"items":384},true,[385,387,389,391,393,395,397,399],[386],"Build the single-agent baseline first, with the same tools and a fair token budget.",[388],"Write down the isolation argument: what stays out of whose context.",[390],"Classify the work as read-heavy or write-heavy. Keep writes single-threaded.",[392],"Give every worker a full brief: objective, output format, tools, boundaries and a stop condition, since it inherits nothing.",[394],"Put explicit scaling rules in the lead prompt and a hard cap on workers and turns.",[396],"Add a verification step on the merged result, and decide who is allowed to say \"done\".",[398],"Evaluate both versions on about 20 realistic queries first, as Anthropic suggests, then grow the set. Compare quality, tokens, latency and failure rate.",[400],"Log every delegation with its brief and its summary so you can debug lost context, and plan how running agents survive a deployment.",{"type":131,"content":402},[403,404,408],"On the last point, Anthropic notes that agent processes are stateful and long-running, so it uses rainbow deployments that shift traffic gradually instead of interrupting runs in progress. Cost control is part of the same discipline; the tools in ",{"tag":146,"to":405,"children":406},"\u002Fblog\u002Ftoken-saving-tools-coding-agents-top-20",[407],"token-saving tools for coding agents"," apply to workers too.",{"type":152,"level":153,"id":124,"text":125},{"type":131,"content":411},[412],"For a typical company use case, I would ship a single agent with strong tools and a good harness, and add exactly one multi-agent element where the isolation argument is strong: a research subagent that returns cited summaries, or a clean-context reviewer on the output. Both keep the writes with one agent, and both are easy to measure against the baseline.",{"type":131,"content":414},[415,416,231],"I would avoid swarms of peers, open-ended debate between same-model agents, and parallel code writers until the evidence changes. Even Cognition, the loudest critic, now describes a narrower class of multi-agent designs that work. The recent vendor posts I read describe the same shape: one orchestrator that owns the context, with isolated helpers that return summaries. That is the architecture worth learning, and it is a good fit for the kind of ",{"tag":146,"to":417,"children":418},"\u002Fexpertise\u002Fai-engineer",[419],"AI engineering work I do for clients",{"type":152,"level":153,"id":127,"text":128},{"type":170,"ordered":383,"items":422},[423,426,429,432,435,438,441,444,447,450,453],[424],{"tag":179,"href":37,"children":425},[36],[427],{"tag":179,"href":40,"children":428},[39],[430],{"tag":179,"href":43,"children":431},[42],[433],{"tag":179,"href":46,"children":434},[45],[436],{"tag":179,"href":49,"children":437},[48],[439],{"tag":179,"href":52,"children":440},[51],[442],{"tag":179,"href":55,"children":443},[54],[445],{"tag":179,"href":58,"children":446},[57],[448],{"tag":179,"href":61,"children":449},[60],[451],{"tag":179,"href":64,"children":452},[63],[454],{"tag":179,"href":67,"children":455},[66],[457,525,586,664],{"slug":458,"published":5,"minutes":459,"category":7,"tags":460,"keywords":464,"about":474,"sources":482,"cover":519,"og":520,"expertise":70,"locales":521,"lang":72,"title":522,"description":523,"coverAlt":524},"openai-dots-always-on-agents-impact",14,[461,10,462,463],"OpenAI dots","GPT-6 Astra","AI governance",[461,465,466,467,468,469,470,471,472,473],"what are OpenAI dots","OpenAI dots impact","always-on AI agents","GPT-6 Astra agents","dots Auto-review and Custom Rules","specialist dots for enterprise","OpenAI dots EU availability","OpenAI DevDay 2026","AI agents in the workplace",[475,478,481],{"name":476,"url":477},"OpenAI","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FOpenAI",{"name":479,"url":480},"ChatGPT","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FChatGPT",{"name":29,"url":30},[483,486,489,492,495,498,501,504,507,510,513,516],{"title":484,"url":485},"OpenAI: Introducing dots (29 September 2026)","https:\u002F\u002Fopenai.com\u002Findex\u002Fintroducing-dots\u002F",{"title":487,"url":488},"OpenAI: How we build safety, security and privacy into dots","https:\u002F\u002Fopenai.com\u002Findex\u002Fhow-we-build-safety-security-and-privacy-into-dots\u002F",{"title":490,"url":491},"OpenAI Help Center: Dots privacy, security, and safety FAQs","https:\u002F\u002Fhelp.openai.com\u002Fen\u002Farticles\u002F20001529-dots-privacy-security-and-safety-faqs",{"title":493,"url":494},"TechCrunch: OpenAI launches Dots, its bubbly agentic avatar","https:\u002F\u002Ftechcrunch.com\u002F2026\u002F09\u002F29\u002Fopenai-launches-dots-its-bubbly-agentic-avatar\u002F",{"title":496,"url":497},"Unite.AI: OpenAI rolls out dots agents powered by GPT-6 Astra in ChatGPT","https:\u002F\u002Fwww.unite.ai\u002Fopenai-rolls-out-dots-agents-powered-by-gpt-6-astra-in-chatgpt\u002F",{"title":499,"url":500},"MediaNama: OpenAI launches dots that keep working without user prompts","https:\u002F\u002Fwww.medianama.com\u002F2026\u002F10\u002F223-openai-launches-dots-devday-2026\u002F",{"title":502,"url":503},"PYMNTS: OpenAI launches dots to capture AI agent market","https:\u002F\u002Fwww.pymnts.com\u002Fnews\u002Fartificial-intelligence\u002F2026\u002Fopenai-launches-dots-to-capture-ai-agent-market\u002F",{"title":505,"url":506},"Yahoo Finance: OpenAI debuts Dots AI agents in challenge to Meta's Muse","https:\u002F\u002Ffinance.yahoo.com\u002Ftechnology\u002Farticle\u002Fopenai-debuts-dots-ai-agents-in-challenge-to-metas-popular-muse-agent-174616593.html",{"title":508,"url":509},"CNBC: OpenAI abandons plan to release upcoming model as safety concerns escalate","https:\u002F\u002Fwww.cnbc.com\u002F2026\u002F09\u002F28\u002Fopenai-abandons-plan-to-release-upcoming-model-as-safety-concerns-escalate.html",{"title":511,"url":512},"The Hacker News: OpenAI shelves GPT-6.1 Astra after tests find deception and unauthorized actions","https:\u002F\u002Fthehackernews.com\u002F2026\u002F09\u002Fopenai-shelves-gpt-61-astra-after-tests.html",{"title":514,"url":515},"Al Jazeera: OpenAI launches dots, personal AI assistant built to handle everything","https:\u002F\u002Fwww.aljazeera.com\u002Feconomy\u002F2026\u002F9\u002F30\u002Fopenai-launches-dots-personal-ai-assistant-built-to-handle-everything",{"title":517,"url":518},"RedactSure: Do OpenAI dots Custom Rules control what the agent sees?","https:\u002F\u002Fredactsure.com\u002Fresearch\u002Fdo-openai-dots-custom-rules-control-what-the-agent-sees","\u002Fimages\u002Fblog\u002Fopenai-dots-always-on-agents-impact\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fopenai-dots-always-on-agents-impact\u002Fog.jpg",[72,73,74],"OpenAI dots: what always-on agents will change, and what they will not","OpenAI dots are always-on GPT-6 Astra agents with their own computer. What launched, how the safeguards work, and what changes for work, IT, SaaS and Europe.","Diagram: a dot running on GPT-6 Astra fans out to Slack and Teams, more than 4,000 apps, its own cloud computer, and a person who approves and reviews.",{"slug":526,"published":5,"minutes":527,"category":7,"tags":528,"keywords":532,"about":543,"sources":551,"cover":580,"og":581,"expertise":70,"locales":582,"lang":72,"title":583,"description":584,"coverAlt":585},"human-in-the-loop-ai-agents",13,[529,10,530,531],"Human in the loop","Approval gates","Agent safety",[533,534,535,536,537,538,539,540,541,542],"human in the loop AI agents","human in the loop KI-Agent","Mensch im Loop KI","AI agent approval gates","agent approval fatigue","LangGraph interrupt human in the loop","OpenAI Agents SDK needs_approval","Claude Agent SDK canUseTool","AI agent audit trail","risk tiers for AI agent actions",[544,547,548],{"name":545,"url":546},"Human-in-the-loop","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FHuman-in-the-loop",{"name":29,"url":30},{"name":549,"url":550},"Audit trail","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FAudit_trail",[552,555,558,561,564,567,570,573,576,579],{"title":553,"url":554},"LangChain docs: LangGraph interrupts","https:\u002F\u002Fdocs.langchain.com\u002Foss\u002Fpython\u002Flanggraph\u002Finterrupts",{"title":556,"url":557},"LangChain docs: Human-in-the-loop (HumanInTheLoopMiddleware)","https:\u002F\u002Fdocs.langchain.com\u002Foss\u002Fpython\u002Flangchain\u002Fhuman-in-the-loop",{"title":559,"url":560},"OpenAI Agents SDK (Python): Human in the loop","https:\u002F\u002Fgithub.com\u002Fopenai\u002Fopenai-agents-python\u002Fblob\u002Fmain\u002Fdocs\u002Fhuman_in_the_loop.md",{"title":562,"url":563},"Claude Agent SDK: Handle approvals and user input","https:\u002F\u002Fcode.claude.com\u002Fdocs\u002Fen\u002Fagent-sdk\u002Fuser-input",{"title":565,"url":566},"Claude Agent SDK: Configure permissions","https:\u002F\u002Fcode.claude.com\u002Fdocs\u002Fen\u002Fagent-sdk\u002Fpermissions",{"title":568,"url":569},"Claude Code docs: Hooks (PreToolUse defer, PermissionRequest)","https:\u002F\u002Fcode.claude.com\u002Fdocs\u002Fen\u002Fhooks",{"title":571,"url":572},"Anthropic Engineering: Claude Code auto mode","https:\u002F\u002Fanthropic.com\u002Fengineering\u002Fclaude-code-auto-mode",{"title":574,"url":575},"DevOps.com: Anthropic makes Claude Code auto mode the default","https:\u002F\u002Fdevops.com\u002Fanthropic-makes-claude-codes-auto-mode-the-default-betting-automation-beats-manual-review\u002F",{"title":577,"url":578},"EU AI Act, Article 14: Human oversight","https:\u002F\u002Fartificialintelligenceact.eu\u002Farticle\u002F14\u002F",{"title":487,"url":488},"\u002Fimages\u002Fblog\u002Fhuman-in-the-loop-ai-agents\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fhuman-in-the-loop-ai-agents\u002Fog.jpg",[72,73,74],"Human in the loop for AI agents: where to put approval gates","Where approval gates belong in an AI agent, how to avoid rubber-stamping, and how interrupt and resume work in LangGraph and the OpenAI and Claude agent SDKs.","Diagram: an agent proposes an action, a risk gate sends it to automatic execution, to a human approval, or to a block, and every decision lands in an audit log.",{"slug":587,"published":5,"minutes":527,"category":7,"tags":588,"keywords":592,"about":603,"sources":611,"cover":658,"og":659,"expertise":70,"locales":660,"lang":72,"title":661,"description":662,"coverAlt":663},"ai-agent-memory-design",[589,12,590,591],"AI agent memory","Memory poisoning","GDPR",[593,594,595,596,597,598,599,600,601,602],"AI agent memory design","long-term memory for AI agents","episodic semantic procedural memory LLM","agent memory architecture","ChatGPT memory vs Claude memory","Claude memory tool","AI memory poisoning","LLM context compaction","AI agent memory GDPR","short-term vs long-term memory agents",[604,605,608],{"name":29,"url":30},{"name":606,"url":607},"General Data Protection Regulation","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FGeneral_Data_Protection_Regulation",{"name":609,"url":610},"Prompt injection","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FPrompt_injection",[612,615,618,621,624,627,630,633,636,639,642,645,648,649,652,655],{"title":613,"url":614},"Anthropic docs: Memory tool","https:\u002F\u002Fplatform.claude.com\u002Fdocs\u002Fen\u002Fagents-and-tools\u002Ftool-use\u002Fmemory-tool",{"title":616,"url":617},"Anthropic docs: Context editing","https:\u002F\u002Fplatform.claude.com\u002Fdocs\u002Fen\u002Fbuild-with-claude\u002Fcontext-editing",{"title":619,"url":620},"Anthropic Engineering: Effective context engineering for AI agents","https:\u002F\u002Fwww.anthropic.com\u002Fengineering\u002Feffective-context-engineering-for-ai-agents",{"title":622,"url":623},"Sumers et al.: Cognitive Architectures for Language Agents (CoALA)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2309.02427",{"title":625,"url":626},"Packer et al.: MemGPT, Towards LLMs as Operating Systems","https:\u002F\u002Farxiv.org\u002Fabs\u002F2310.08560",{"title":628,"url":629},"Park et al.: Generative Agents, Interactive Simulacra of Human Behavior","https:\u002F\u002Farxiv.org\u002Fabs\u002F2304.03442",{"title":631,"url":632},"Unit 42: When AI Remembers Too Much, persistent behaviors in agents memory","https:\u002F\u002Funit42.paloaltonetworks.com\u002Findirect-prompt-injection-poisons-ai-longterm-memory\u002F",{"title":634,"url":635},"From Untrusted Input to Trusted Memory: A Systematic Study of Memory Poisoning Attacks in LLM Agents (preprint)","https:\u002F\u002Farxiv.org\u002Fhtml\u002F2606.04329v1",{"title":637,"url":638},"The Hacker News: ChatGPT macOS flaw could have enabled long-term spyware via memory function","https:\u002F\u002Fthehackernews.com\u002F2024\u002F09\u002Fchatgpt-macos-flaw-couldve-enabled-long.html",{"title":640,"url":641},"Vectorize: OWASP ASI06, Memory and Context Poisoning explained","https:\u002F\u002Fvectorize.io\u002Farticles\u002Fowasp-asi06",{"title":643,"url":644},"Claude Help Center: Use chat search and memory to build on previous context","https:\u002F\u002Fsupport.claude.com\u002Fen\u002Farticles\u002F11817273-use-claude-s-chat-search-and-memory-to-build-on-previous-context",{"title":646,"url":647},"OpenAI Help Center: Memory in ChatGPT","https:\u002F\u002Fhelp.openai.com\u002Fen\u002Farticles\u002F8590148-memory-faq",{"title":490,"url":491},{"title":650,"url":651},"Flavio Copes: A deep dive into OpenAI dots (quotes the dots documentation on memory)","https:\u002F\u002Fflaviocopes.com\u002Fopenai-dots\u002F",{"title":653,"url":654},"GDPR Article 5: Principles relating to processing of personal data","https:\u002F\u002Fgdpr-info.eu\u002Fart-5-gdpr\u002F",{"title":656,"url":657},"GDPR Article 17: Right to erasure","https:\u002F\u002Fgdpr-info.eu\u002Fart-17-gdpr\u002F","\u002Fimages\u002Fblog\u002Fai-agent-memory-design\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fai-agent-memory-design\u002Fog.jpg",[72,73,74],"Designing memory for AI agents: tiers, write rules, poisoning and GDPR","How to design AI agent memory: context vs session vs long-term tiers, what to write and never store, retrieval, compaction, poisoning and GDPR erasure.","Diagram: nested memory layers of an AI agent, from the working context window through session state to long-term episodic and semantic memory.",{"slug":665,"published":5,"minutes":527,"category":7,"tags":666,"keywords":672,"about":683,"sources":693,"cover":748,"og":749,"expertise":70,"locales":750,"lang":72,"title":751,"description":752,"coverAlt":753},"voice-agents-realtime-latency",[667,668,669,670,671],"Voice agents","Realtime API","Latency","Telephony","AI Act",[673,674,675,676,677,678,679,680,681,682],"voice agents","speech-to-speech vs STT LLM TTS","OpenAI Realtime API","Gemini Live API","voice agent latency budget","turn detection and barge-in","Pipecat vs LiveKit","AI voice agent SIP telephony","German Hungarian voice AI","AI Act Article 50 voice bot disclosure",[684,687,690],{"name":685,"url":686},"Voice user interface","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FVoice_user_interface",{"name":688,"url":689},"Speech recognition","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FSpeech_recognition",{"name":691,"url":692},"Artificial Intelligence Act","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FArtificial_Intelligence_Act",[694,697,700,703,706,709,712,715,718,721,724,727,730,733,736,739,742,745],{"title":695,"url":696},"OpenAI: Voice agents guide","https:\u002F\u002Fdevelopers.openai.com\u002Fapi\u002Fdocs\u002Fguides\u002Fvoice-agents",{"title":698,"url":699},"OpenAI: gpt-realtime model page","https:\u002F\u002Fdevelopers.openai.com\u002Fapi\u002Fdocs\u002Fmodels\u002Fgpt-realtime",{"title":701,"url":702},"OpenAI: Voice activity detection in the Realtime API","https:\u002F\u002Fdevelopers.openai.com\u002Fapi\u002Fdocs\u002Fguides\u002Frealtime-vad",{"title":704,"url":705},"OpenAI: Realtime API with SIP","https:\u002F\u002Fdevelopers.openai.com\u002Fapi\u002Fdocs\u002Fguides\u002Frealtime-sip",{"title":707,"url":708},"OpenAI: Realtime conversations (function calling, interruption)","https:\u002F\u002Fdevelopers.openai.com\u002Fapi\u002Fdocs\u002Fguides\u002Frealtime-conversations",{"title":710,"url":711},"Google: Gemini Live API overview","https:\u002F\u002Fai.google.dev\u002Fgemini-api\u002Fdocs\u002Flive",{"title":713,"url":714},"Google: Gemini Live API capabilities guide","https:\u002F\u002Fai.google.dev\u002Fgemini-api\u002Fdocs\u002Flive-guide",{"title":716,"url":717},"Deepgram: Flux quickstart","https:\u002F\u002Fdevelopers.deepgram.com\u002Fdocs\u002Fflux\u002Fquickstart",{"title":719,"url":720},"Deepgram: Models and languages overview","https:\u002F\u002Fdevelopers.deepgram.com\u002Fdocs\u002Fmodels-languages-overview",{"title":722,"url":723},"ElevenLabs: Agents platform overview","https:\u002F\u002Felevenlabs.io\u002Fdocs\u002Feleven-agents\u002Foverview",{"title":725,"url":726},"ElevenLabs: Text to speech models and languages","https:\u002F\u002Felevenlabs.io\u002Fdocs\u002Foverview\u002Fcapabilities\u002Ftext-to-speech",{"title":728,"url":729},"LiveKit: Agents overview","https:\u002F\u002Fdocs.livekit.io\u002Fagents\u002F",{"title":731,"url":732},"LiveKit: Turn detector","https:\u002F\u002Fdocs.livekit.io\u002Fagents\u002Flogic\u002Fturns\u002Fturn-detector\u002F",{"title":734,"url":735},"Pipecat: Introduction","https:\u002F\u002Fdocs.pipecat.ai\u002Fgetting-started\u002Fintroduction",{"title":737,"url":738},"Pipecat: Smart Turn model (GitHub)","https:\u002F\u002Fgithub.com\u002Fpipecat-ai\u002Fsmart-turn",{"title":740,"url":741},"Fora Soft: Voice AI agents on LiveKit, 2026 engineer playbook","https:\u002F\u002Fwww.forasoft.com\u002Fblog\u002Farticle\u002Fvoice-ai-agents-livekit-guide",{"title":743,"url":744},"EU AI Act: Article 50, transparency obligations","https:\u002F\u002Fartificialintelligenceact.eu\u002Farticle\u002F50\u002F",{"title":746,"url":747},"Jones Walker: Yes, August 2 still matters (AI Act delay and Article 50)","https:\u002F\u002Fwww.joneswalker.com\u002Fen\u002Finsights\u002Fblogs\u002Fai-law-blog\u002Fyes-august-2-still-matters-the-eu-approved-a-high-risk-ai-delay-but-most-trans.html?id=102nbon","\u002Fimages\u002Fblog\u002Fvoice-agents-realtime-latency\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fvoice-agents-realtime-latency\u002Fog.jpg",[72,73,74],"Building voice agents: realtime speech-to-speech or STT, LLM and TTS?","Realtime speech-to-speech or a cascaded pipeline? Latency budget per stage, turn-taking, tool calls, SIP, German and Hungarian quality, and AI Act disclosure.","Diagram: a caller reaches a voice agent over SIP or WebRTC, which fans out to turn detection, speech recognition, an LLM with tools and speech synthesis.",1791009037088]