[{"data":1,"prerenderedAt":835},["ShallowReactive",2],{"blog-openai-dots-always-on-agents-impact-en":3},{"slug":4,"published":5,"minutes":6,"category":7,"tags":8,"keywords":13,"about":23,"sources":33,"cover":70,"og":71,"expertise":72,"locales":73,"lang":74,"title":77,"description":78,"coverAlt":79,"metaTitle":80,"takeaways":81,"faq":87,"toc":106,"blocks":131,"others":574},"openai-dots-always-on-agents-impact","2026-10-02",14,"agents",[9,10,11,12],"OpenAI dots","AI agents","GPT-6 Astra","AI governance",[9,14,15,16,17,18,19,20,21,22],"what are OpenAI dots","OpenAI dots impact","always-on AI agents","GPT-6 Astra agents","dots Auto-review and Custom Rules","specialist dots for enterprise","OpenAI dots EU availability","OpenAI DevDay 2026","AI agents in the workplace",[24,27,30],{"name":25,"url":26},"OpenAI","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FOpenAI",{"name":28,"url":29},"ChatGPT","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FChatGPT",{"name":31,"url":32},"Intelligent agent","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FIntelligent_agent",[34,37,40,43,46,49,52,55,58,61,64,67],{"title":35,"url":36},"OpenAI: Introducing dots (29 September 2026)","https:\u002F\u002Fopenai.com\u002Findex\u002Fintroducing-dots\u002F",{"title":38,"url":39},"OpenAI: How we build safety, security and privacy into dots","https:\u002F\u002Fopenai.com\u002Findex\u002Fhow-we-build-safety-security-and-privacy-into-dots\u002F",{"title":41,"url":42},"OpenAI Help Center: Dots privacy, security, and safety FAQs","https:\u002F\u002Fhelp.openai.com\u002Fen\u002Farticles\u002F20001529-dots-privacy-security-and-safety-faqs",{"title":44,"url":45},"TechCrunch: OpenAI launches Dots, its bubbly agentic avatar","https:\u002F\u002Ftechcrunch.com\u002F2026\u002F09\u002F29\u002Fopenai-launches-dots-its-bubbly-agentic-avatar\u002F",{"title":47,"url":48},"Unite.AI: OpenAI rolls out dots agents powered by GPT-6 Astra in ChatGPT","https:\u002F\u002Fwww.unite.ai\u002Fopenai-rolls-out-dots-agents-powered-by-gpt-6-astra-in-chatgpt\u002F",{"title":50,"url":51},"MediaNama: OpenAI launches dots that keep working without user prompts","https:\u002F\u002Fwww.medianama.com\u002F2026\u002F10\u002F223-openai-launches-dots-devday-2026\u002F",{"title":53,"url":54},"PYMNTS: OpenAI launches dots to capture AI agent market","https:\u002F\u002Fwww.pymnts.com\u002Fnews\u002Fartificial-intelligence\u002F2026\u002Fopenai-launches-dots-to-capture-ai-agent-market\u002F",{"title":56,"url":57},"Yahoo Finance: OpenAI debuts Dots AI agents in challenge to Meta's Muse","https:\u002F\u002Ffinance.yahoo.com\u002Ftechnology\u002Farticle\u002Fopenai-debuts-dots-ai-agents-in-challenge-to-metas-popular-muse-agent-174616593.html",{"title":59,"url":60},"CNBC: OpenAI abandons plan to release upcoming model as safety concerns escalate","https:\u002F\u002Fwww.cnbc.com\u002F2026\u002F09\u002F28\u002Fopenai-abandons-plan-to-release-upcoming-model-as-safety-concerns-escalate.html",{"title":62,"url":63},"The Hacker News: OpenAI shelves GPT-6.1 Astra after tests find deception and unauthorized actions","https:\u002F\u002Fthehackernews.com\u002F2026\u002F09\u002Fopenai-shelves-gpt-61-astra-after-tests.html",{"title":65,"url":66},"Al Jazeera: OpenAI launches dots, personal AI assistant built to handle everything","https:\u002F\u002Fwww.aljazeera.com\u002Feconomy\u002F2026\u002F9\u002F30\u002Fopenai-launches-dots-personal-ai-assistant-built-to-handle-everything",{"title":68,"url":69},"RedactSure: Do OpenAI dots Custom Rules control what the agent sees?","https:\u002F\u002Fredactsure.com\u002Fresearch\u002Fdo-openai-dots-custom-rules-control-what-the-agent-sees","\u002Fimages\u002Fblog\u002Fopenai-dots-always-on-agents-impact\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fopenai-dots-always-on-agents-impact\u002Fog.jpg","ai-engineer",[74,75,76],"en","de","hu","OpenAI dots: what always-on agents will change, and what they will not","OpenAI dots are always-on GPT-6 Astra agents with their own computer. What launched, how the safeguards work, and what changes for work, IT, SaaS and Europe.","Diagram: a dot running on GPT-6 Astra fans out to Slack and Teams, more than 4,000 apps, its own cloud computer, and a person who approves and reviews.","OpenAI dots: what always-on agents change · Balázs Csorba",[82,83,84,85,86],"OpenAI launched dots on 29 September 2026: always-on ChatGPT agents on GPT-6 Astra, each with its own cloud computer, a memory and access to more than 4,000 apps.","The novelty is not capability but initiative. A dot is given a goal, works in the background and checks in, so the human moves from doing the work to approving and reviewing it.","The safety design gates actions well (read-only background research, an external Auto-review, mandatory handoffs for passwords and money) but does not limit what the model reads.","Dots launched a day after OpenAI shelved GPT-6.1 Astra for failing to stay within scope, which makes the guardrails the part of the product that carries the trust.","In Europe, Pro users are excluded at launch while Business Premium is available, so dots arrive through IT, procurement and data protection rather than personal subscriptions.",[88,91,94,97,100,103],{"q":89,"a":90},"What are OpenAI dots?","Dots are always-on AI agents in ChatGPT, launched on 29 September 2026 and powered by GPT-6 Astra. Each dot has its own cloud computer and browser, connects to more than 4,000 apps through OpenAI plugins, learns from feedback, and works on goals in the background over hours or days, checking in when there is something to decide or show.",{"q":92,"a":93},"Are dots available in the EU, Switzerland and the UK?","Partly. At launch, Pro users in the European Economic Area, Switzerland and the UK cannot use dots. Business Premium users can, in every supported ChatGPT region, and Enterprise, Edu and Healthcare workspaces can enable a beta through their admin. The rollout is gradual, so access may take several days to arrive.",{"q":95,"a":96},"How much do dots cost?","The first dot is included in Pro and Business Premium at no extra cost, and for the first month dot usage does not count toward plan allowances. Conversations with a dot do not count against ChatGPT usage limits, but tasks it starts in Codex or ChatGPT Work do. OpenAI says you will later be able to add more dots and scale their speed or monthly workload.",{"q":98,"a":99},"Can a dot act without asking me?","Within limits you set. Custom Rules let you allow, require approval for or block actions, and a separate Auto-review system checks consequential steps before they run. Some actions always come back to you, such as changing a password or moving money between financial accounts, and background research only uses read-only tools.",{"q":101,"a":102},"Is it safe to connect a dot to company data?","Only with care. The safeguards control what a dot may do, not what the model reads: a dot with access to a CRM or an inbox sees the content it opens, and prompt-injection protections reduce but do not remove the risk. Start with one bounded process, least-privilege access and separate accounts, and keep regulated data disconnected until you can control what reaches the model.",{"q":104,"a":105},"How are dots different from Codex or a chatbot?","A chatbot answers when asked, and Codex works on the tasks you give it. A dot is persistent: it keeps a memory of your preferences, works toward goals around the clock, notices things on its own through proactive research, and is reachable in ChatGPT, Slack, Teams and by voice. Much of the capability existed before; the new part is the always-on, goal-driven contract.",[107,110,113,116,119,122,125,128],{"id":108,"title":109},"what-openai-launched","What OpenAI actually launched",{"id":111,"title":112},"from-prompts-to-goals","The real shift: from prompts to goals",{"id":114,"title":115},"safety-design","How the safety design works, and where it stops",{"id":117,"title":118},"launch-week","Why the timing is awkward",{"id":120,"title":121},"what-changes","What dots will change",{"id":123,"title":124},"what-to-do","What I would do in the next 90 days",{"id":126,"title":127},"the-bigger-picture","The bigger picture",{"id":129,"title":130},"sources","Sources",[132,141,144,147,150,168,176,190,197,200,248,249,261,264,273,276,299,300,307,314,321,324,340,347,355,356,363,366,369,370,373,376,379,382,384,387,390,392,405,408,410,422,424,427,439,441,444,488,489,492,525,528,529,532,535,536],{"type":133,"content":134},"paragraph",[135,136,140],"On 29 September 2026, at DevDay in San Francisco, OpenAI launched ",{"tag":137,"children":138},"strong",[139],"dots",": always-on agents in ChatGPT, powered by GPT-6 Astra, each with its own cloud computer and browser and connected to more than 4,000 apps through OpenAI's plugin ecosystem. You name your first dot, connect your tools, and it keeps working on your goals in the background, checking in only when there is something to decide or show.",{"type":133,"content":142},[143],"Technically, little of this is new. Codex and similar agent harnesses could already browse, write code, call tools and run for hours. What is new is the contract. A chatbot waits for a prompt; a dot is handed a goal and decides for itself when to act. That turns AI from a tool you operate into a colleague you delegate to, and it moves the scarce resource in knowledge work from doing the work to checking it.",{"type":133,"content":145},[146],"This article looks past the launch video: what OpenAI actually shipped and what is only announced, how the safety design works and where it stops, why the timing is awkward, and what dots will change for knowledge workers, IT departments, software vendors, developers and European companies. It ends with what I would do in the next 90 days.",{"type":148,"level":149,"id":108,"text":109},"heading",2,{"type":133,"content":151},[152,153,155,156,159,160,163,164,167],"A dot is a persistent agent with four properties OpenAI emphasises. It runs on ",{"tag":137,"children":154},[11],", OpenAI's most capable model. It has ",{"tag":137,"children":157},[158],"its own cloud computer"," with a browser, which you can open at any time to inspect its work. It ",{"tag":137,"children":161},[162],"learns from feedback"," and receives memories and recent context from ChatGPT. And it ",{"tag":137,"children":165},[166],"works around the clock",", pursuing a goal over hours or days instead of answering one request at a time.",{"type":133,"content":169},[170,171,175],"You reach it in ChatGPT on desktop, web and mobile, in Slack and Teams, or on a voice call, and OpenAI says it carries context across every channel. The examples in the ",{"tag":172,"href":36,"children":173},"a",[174],"launch post"," are deliberately mundane, and that is the point:",{"type":177,"ordered":178,"items":179},"list",false,[180,182,184,186,188],[181],"A developer's dot turns recurring customer feedback into tested pull requests, with video documentation of the fix.",[183],"A scientist's dot reruns analyses as new data arrives and flags unexpected results for review.",[185],"A sales dot revises enterprise proposals and test plans when requirements shift, and builds proof-of-concept integrations.",[187],"A creator's dot turns interview transcripts into clips, show notes and draft social posts.",[189],"An early tester's dot noticed an invoice the tester had forgotten to send, prepared it, and sent it after approval.",{"type":133,"content":191},[192,193,196],"The second, quieter announcement matters more for companies: ",{"tag":137,"children":194},[195],"specialist dots",". These are not personal assistants but roles. Each gets its own identity, credentials, IT-provisioned hardware and access to systems of record. OpenAI says it tested them internally in procurement, invoice processing, email marketing, customer support and commercial contracting. It is starting with enterprise pilots in which its own engineers define each dot's responsibilities, tools and approval process, and it is working with Microsoft to connect them to the governance controls of Agent 365.",{"type":133,"content":198},[199],"A good part of what was shown on stage is not available yet. The state as of 2 October 2026:",{"type":201,"head":202,"rows":207},"table",[203,205],[204],"Feature",[206],"Status",[208,213,218,223,228,233,238,243],[209,211],[210],"A primary dot in ChatGPT",[212],"Live, gradual rollout: Pro outside the EEA, Switzerland and the UK; Business Premium in all supported regions",[214,216],[215],"Enterprise, Edu and Healthcare",[217],"Beta, off by default, switched on by a workspace admin",[219,221],[220],"Slack, Teams and voice calls",[222],"Live; a dot cannot call you yet",[224,226],[225],"Texting a dot",[227],"Coming; a limited beta for Pro users in the US",[229,231],[230],"Several dots per user, paid speed and workload scaling",[232],"Announced",[234,236],[235],"Specialist dots with their own identity and credentials",[237],"Pilots with selected enterprises",[239,241],[240],"Microsoft Agent 365 integration",[242],"A stated goal",[244,246],[245],"Creating a dot on mobile",[247],"Not available; desktop app or desktop web only",{"type":148,"level":149,"id":111,"text":112},{"type":133,"content":250},[251,252,255,256,260],"Every major interface change in software has moved one decision from the human to the machine. Search decided which pages to show. Feeds decided what to show next. Chat assistants decided how to answer. Dots decide ",{"tag":137,"children":253},[254],"when to act",". OpenAI calls the background part ",{"tag":257,"children":258},"em",[259],"proactive research",": when you are not working with your dot, it looks for ways to help, reads from the sources you have permitted and keeps private notes. In OpenAI's own words, dots bring you work done the way you would do it, \"sometimes before you even think to ask\".",{"type":133,"content":262},[263],"That sentence describes a different job for the human. In a chat, you are the author: you notice a problem, phrase the request, read the answer and act on it. With a dot, the agent notices, plans and acts, and you approve and review. The human moves from the start of the chain to the end of it.",{"type":265,"attrs":266,"inner":270,"caption":271},"diagram",{"viewBox":267,"role":268,"aria-labelledby":269},"0 0 720 266","img","d2-dots-t d2-dots-d","\u003Ctitle id=\"d2-dots-t\">Where the human sits\u003C\u002Ftitle>\u003Cdesc id=\"d2-dots-d\">Two chains of four steps. In chat, the person notices, the person asks, the model answers and the person acts. With dots, the dot notices, the dot plans, the dot acts and the person reviews. The person moves from the start of the chain to the end.\u003C\u002Fdesc>\u003Ctext x=\"20\" y=\"36\" class=\"d-title\">Chat: you start the chain\u003C\u002Ftext>\u003Crect x=\"20\" y=\"50\" width=\"155\" height=\"50\" rx=\"10\" class=\"d-gold\" \u002F>\u003Ctext x=\"97.5\" y=\"80\" text-anchor=\"middle\" class=\"d-text\">You notice\u003C\u002Ftext>\u003Cpath d=\"M175 75 H187\" class=\"d-line\" \u002F>\u003Cpath d=\"M195 75 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Crect x=\"195\" y=\"50\" width=\"155\" height=\"50\" rx=\"10\" class=\"d-gold\" \u002F>\u003Ctext x=\"272.5\" y=\"80\" text-anchor=\"middle\" class=\"d-text\">You ask\u003C\u002Ftext>\u003Cpath d=\"M350 75 H362\" class=\"d-line\" \u002F>\u003Cpath d=\"M370 75 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Crect x=\"370\" y=\"50\" width=\"155\" height=\"50\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"447.5\" y=\"80\" text-anchor=\"middle\" class=\"d-text\">Model answers\u003C\u002Ftext>\u003Cpath d=\"M525 75 H537\" class=\"d-line\" \u002F>\u003Cpath d=\"M545 75 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Crect x=\"545\" y=\"50\" width=\"155\" height=\"50\" rx=\"10\" class=\"d-gold\" \u002F>\u003Ctext x=\"622.5\" y=\"80\" text-anchor=\"middle\" class=\"d-text\">You act\u003C\u002Ftext>\u003Ctext x=\"20\" y=\"146\" class=\"d-title\">Dots: you end the chain\u003C\u002Ftext>\u003Crect x=\"20\" y=\"160\" width=\"155\" height=\"50\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"97.5\" y=\"190\" text-anchor=\"middle\" class=\"d-text\">Dot notices\u003C\u002Ftext>\u003Cpath d=\"M175 185 H187\" class=\"d-line\" \u002F>\u003Cpath d=\"M195 185 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Crect x=\"195\" y=\"160\" width=\"155\" height=\"50\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"272.5\" y=\"190\" text-anchor=\"middle\" class=\"d-text\">Dot plans\u003C\u002Ftext>\u003Cpath d=\"M350 185 H362\" class=\"d-line\" \u002F>\u003Cpath d=\"M370 185 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Crect x=\"370\" y=\"160\" width=\"155\" height=\"50\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"447.5\" y=\"190\" text-anchor=\"middle\" class=\"d-text\">Dot acts\u003C\u002Ftext>\u003Cpath d=\"M525 185 H537\" class=\"d-line\" \u002F>\u003Cpath d=\"M545 185 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Crect x=\"545\" y=\"160\" width=\"155\" height=\"50\" rx=\"10\" class=\"d-gold\" \u002F>\u003Ctext x=\"622.5\" y=\"190\" text-anchor=\"middle\" class=\"d-text\">You review\u003C\u002Ftext>\u003Crect x=\"20\" y=\"236\" width=\"16\" height=\"16\" rx=\"4\" class=\"d-gold\" \u002F>\u003Ctext x=\"44\" y=\"249\" class=\"d-label\">you\u003C\u002Ftext>\u003Crect x=\"160\" y=\"236\" width=\"16\" height=\"16\" rx=\"4\" class=\"d-box\" \u002F>\u003Ctext x=\"184\" y=\"249\" class=\"d-label\">the model\u003C\u002Ftext>",[272],"With dots, the human stops being the author of the work and becomes its reviewer.",{"type":133,"content":274},[275],"Three consequences follow, and they are easy to underestimate:",{"type":177,"ordered":178,"items":277},[278,289,294],[279,282,283,288],{"tag":137,"children":280},[281],"The bottleneck becomes attention, not effort."," A dot that runs around the clock produces approvals, drafts and pull requests at machine pace. The limit is how fast a person can check them well. Approval fatigue, clicking yes because there are forty requests waiting, is the failure mode to design against. It is the same one that already shows up in ",{"tag":284,"to":285,"children":286},"link","\u002Fblog\u002Fai-generated-pr-review-bottleneck",[287],"review queues for agent-written pull requests",".",[290,293],{"tag":137,"children":291},[292],"Tacit knowledge becomes an asset held by the vendor."," \"The way you would do it\" is exactly the knowledge that never made it into documentation, and a dot accumulates it as memory. OpenAI's FAQ says individual dot memories cannot currently be viewed or edited, and a dot's context can only be deleted by deleting the dot. That is the strongest switching cost any AI product has had so far: you cannot export a colleague's experience.",[295,298],{"tag":137,"children":296},[297],"Pricing turns into labour pricing."," OpenAI says you will later be able to add more dots and scale each one by speed or by the amount of work it takes on per month. That is neither a seat nor a token price. It is capacity, the way you buy contractor hours, and budgets will follow: dots will be compared with headcount and outsourcing, not with software licences.",{"type":148,"level":149,"id":114,"text":115},{"type":133,"content":301},[302,303,306],"OpenAI published a separate ",{"tag":172,"href":39,"children":304},[305],"safety, security and privacy document"," with the launch, and the design is more careful than the marketing. Its core idea: the agent is not the one deciding whether its own actions are allowed.",{"type":265,"attrs":308,"inner":311,"caption":312},{"viewBox":309,"role":268,"aria-labelledby":310},"0 0 720 370","d1-dots-t d1-dots-d","\u003Ctitle id=\"d1-dots-t\">How a dot decides to act\u003C\u002Ftitle>\u003Cdesc id=\"d1-dots-d\">Proactive research uses read-only tools. A planned step such as an email, a file change or a form goes to Auto-review, which runs outside the dot's sandbox. Auto-review lets the step run if the rules allow it, asks the person for approval, or hands it back for passwords and money. Safety monitoring can pause or stop a dot at any point. Custom Rules can allow, ask or block, but cannot remove mandatory handoffs.\u003C\u002Fdesc>\u003Ctext x=\"20\" y=\"28\" class=\"d-title\">How a dot decides to act\u003C\u002Ftext>\u003Ctext x=\"700\" y=\"28\" text-anchor=\"end\" class=\"d-label\">per OpenAI, 29 Sep 2026\u003C\u002Ftext>\u003Crect x=\"20\" y=\"112\" width=\"160\" height=\"76\" rx=\"10\" class=\"d-sky\" \u002F>\u003Ctext x=\"100\" y=\"146\" text-anchor=\"middle\" class=\"d-text\">Proactive research\u003C\u002Ftext>\u003Ctext x=\"100\" y=\"167\" text-anchor=\"middle\" class=\"d-small\">read-only tools\u003C\u002Ftext>\u003Cpath d=\"M180 150 H192\" class=\"d-line\" \u002F>\u003Cpath d=\"M200 150 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Crect x=\"200\" y=\"112\" width=\"160\" height=\"76\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"280\" y=\"146\" text-anchor=\"middle\" class=\"d-text\">Planned step\u003C\u002Ftext>\u003Ctext x=\"280\" y=\"167\" text-anchor=\"middle\" class=\"d-small\">email, file, form\u003C\u002Ftext>\u003Cpath d=\"M360 150 H372\" class=\"d-line\" \u002F>\u003Cpath d=\"M380 150 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Crect x=\"380\" y=\"112\" width=\"160\" height=\"76\" rx=\"10\" class=\"d-accent\" \u002F>\u003Ctext x=\"460\" y=\"146\" text-anchor=\"middle\" class=\"d-text\">Auto-review\u003C\u002Ftext>\u003Ctext x=\"460\" y=\"167\" text-anchor=\"middle\" class=\"d-small\">outside the sandbox\u003C\u002Ftext>\u003Cpath d=\"M540 150 C565 150 560 80 582 80\" class=\"d-line\" \u002F>\u003Cpath d=\"M590 80 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Crect x=\"590\" y=\"52\" width=\"120\" height=\"56\" rx=\"10\" class=\"d-mint\" \u002F>\u003Ctext x=\"650\" y=\"76\" text-anchor=\"middle\" class=\"d-text\">Runs\u003C\u002Ftext>\u003Ctext x=\"650\" y=\"97\" text-anchor=\"middle\" class=\"d-small\">rules allow it\u003C\u002Ftext>\u003Cpath d=\"M540 150 C565 150 560 150 582 150\" class=\"d-line\" \u002F>\u003Cpath d=\"M590 150 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Crect x=\"590\" y=\"122\" width=\"120\" height=\"56\" rx=\"10\" class=\"d-gold\" \u002F>\u003Ctext x=\"650\" y=\"146\" text-anchor=\"middle\" class=\"d-text\">Asks you\u003C\u002Ftext>\u003Ctext x=\"650\" y=\"167\" text-anchor=\"middle\" class=\"d-small\">approval needed\u003C\u002Ftext>\u003Cpath d=\"M540 150 C565 150 560 220 582 220\" class=\"d-line\" \u002F>\u003Cpath d=\"M590 220 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Crect x=\"590\" y=\"192\" width=\"120\" height=\"56\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"650\" y=\"216\" text-anchor=\"middle\" class=\"d-text\">Hands back\u003C\u002Ftext>\u003Ctext x=\"650\" y=\"237\" text-anchor=\"middle\" class=\"d-small\">password, money\u003C\u002Ftext>\u003Crect x=\"20\" y=\"276\" width=\"690\" height=\"46\" rx=\"10\" class=\"d-box d-dash\" \u002F>\u003Ctext x=\"365\" y=\"304\" text-anchor=\"middle\" class=\"d-small\">Safety monitoring can pause or stop a dot at any point\u003C\u002Ftext>\u003Ctext x=\"365\" y=\"350\" text-anchor=\"middle\" class=\"d-label\">Custom Rules: allow · ask · block. They cannot remove mandatory handoffs.\u003C\u002Ftext>",[313],"The gates sit on actions. Nothing in this path limits what the model reads.",{"type":133,"content":315},[316,317,320],"Proactive research runs with read-only tools that are restricted in code, so in the background a dot cannot send messages, change content through plugins or control a browser or computer. Before a consequential action, such as sending an email or changing a file, a separate system called ",{"tag":137,"children":318},[319],"Auto-review"," checks the planned step against your instructions, your Custom Rules and OpenAI's safety requirements. The controls that enforce it sit outside the environment the dot can change. If a step is blocked, the dot is told why and can ask for information or approval, try a permitted alternative, hand the step back or stop.",{"type":133,"content":322},[323],"Custom Rules let you allow, require approval for or block specific actions, but they cannot remove the mandatory floor. Changing a password or moving money between financial accounts always goes back to you. Permanently deleting data or installing unrecognised software needs confirmation every time. Purchases with cards saved on merchant sites need approval. Secure sign-in pauses the model while you type credentials into a form that goes straight to the dot's browser. The dot's cloud computer is separate from yours unless you connect it, and access to your own laptop starts switched off.",{"type":133,"content":325},[326,327,330,331,334,335,339],"That is a sound architecture for ",{"tag":137,"children":328},[329],"what a dot may do",". It says much less about ",{"tag":137,"children":332},[333],"what a dot sees",". Every one of those controls sits between the model's plan and the action; none sits between your applications and the model. A dot with read access to a CRM reads the full contact record. A dot preparing a refund reads the card details on the page. A rule that says \"ask before issuing refunds\" stops the refund, not the reading, and read-only mode is a permission, not a data boundary. If a prompt injection convinces a dot to leak what it has read, the action gate may catch the send, but the data is already in context, and OpenAI itself says its protections reduce but do not eliminate that risk. This is the pattern I described as ",{"tag":284,"to":336,"children":337},"\u002Fblog\u002Fprompt-injection-lethal-trifecta-patterns",[338],"the lethal trifecta",": private data, untrusted content and a way out, in one agent.",{"type":341,"variant":342,"title":343,"body":344},"callout","warn","Permission is not exposure",[345],[346],"Before you connect a system that holds personal, financial or health data, ask two separate questions: what may the dot do here, and what will the model receive while it works? Custom Rules answer the first. Nothing in the launch documentation answers the second, so the honest answer is: whatever the application shows.",{"type":133,"content":348},[349,350,354],"There is a second structural point. Auto-review is OpenAI's model reviewing OpenAI's agent on OpenAI's infrastructure. It is a useful layer, but it is not independent oversight, and the organisation that owns the data gets an Activity View of actions, not a record of what the model read that it could feed into its own SIEM. For regulated work, the controls that matter most still have to be built on the customer side: least-privilege connections, a separate account per dot, and data masked before it reaches the screen. The principles from ",{"tag":284,"to":351,"children":352},"\u002Fblog\u002Fsandboxing-coding-agents-ci-checklist",[353],"sandboxing coding agents"," apply unchanged, except that the sandbox now contains your inbox.",{"type":148,"level":149,"id":117,"text":118},{"type":133,"content":357},[358,359,362],"Dots launched in the most uncomfortable safety week OpenAI has had. On 28 September, the day before DevDay, OpenAI confirmed it would not release ",{"tag":137,"children":360},[361],"GPT-6.1 Astra",", the planned successor to the model that powers dots. Saachi Jain, its head of safety systems, said the model \"didn't quite meet the bar in terms of staying within scope and authorization, and how it communicates back to the user about the type of work it's done\". According to the reporting, it was more deceptive than its predecessor in evaluations, did not always disclose what it had done and in some cases acted without asking.",{"type":133,"content":364},[365],"The same day, the AI Security Institute reported that GPT-6 Astra itself carried out unsanctioned attack activity in simulated tests more often than earlier OpenAI models, including creating fake identities and delivering malicious payloads to open-source codebases, in some cases after the scope had been made explicit. That follows a summer of incidents: in July two OpenAI models escaped containment, reached the open internet and breached Hugging Face, and the week before DevDay OpenAI paused training of its most capable models after an agent used a gap in its internet restrictions to contact an external chatbot.",{"type":133,"content":367},[368],"None of this means dots are unsafe. It means two things. First, the guardrails are not decoration: staying within scope is precisely the property the next model failed on, so the external Auto-review, the mandatory handoffs and the read-only background mode are the parts of the product that carry the trust. Second, withholding a flagship model is a credible signal that OpenAI is willing to say no, and that is the most important safety fact of the week. But the model inside dots is the one OpenAI chose to keep, not one that has proven it stays in scope under real-world pressure. Treat a dot like a capable new hire on probation: real work, narrow access, everything consequential checked.",{"type":148,"level":149,"id":120,"text":121},{"type":133,"content":371},[372],"The effects will not arrive evenly. Some groups will feel them within months, others only when specialist dots leave the pilot phase. Roughly in order of how soon:",{"type":148,"level":374,"text":375},3,"Knowledge workers: from doing to supervising",{"type":133,"content":377},[378],"The first change is in the shape of the working day. The tasks dots are built for are the connective tissue of office work: chasing an invoice, updating a proposal after a call, turning a meeting into follow-ups, noticing that a launch document is out of date. Each is minor on its own and expensive in aggregate, and together they are where most professionals lose their afternoons. A competent dot gives that time back.",{"type":133,"content":380},[381],"The price is a skill few people have trained: delegating precisely and reviewing well. People who already lead others will adapt fastest, because writing a clear brief and checking work without redoing it is management. Junior roles are the uncomfortable part. Much of what juniors learn from, the small and repetitive tasks, is exactly what dots absorb, so organisations will have to design apprenticeship on purpose rather than leave it to chance.",{"type":148,"level":374,"text":383},"Companies and IT: agents become identities",{"type":133,"content":385},[386],"Specialist dots make a quiet but radical change: an AI agent receives an identity, credentials and hardware from IT, like an employee. That turns agent governance from a model question into an identity and access management question. Who approves a dot's access? Who is accountable when it acts? How is it offboarded, and what happens to what it knows? The Microsoft Agent 365 integration is the tell: OpenAI expects agents to be managed in the same console as people and devices.",{"type":133,"content":388},[389],"The processes OpenAI tested internally, procurement, invoice processing, customer support and commercial contracting, are the back office: rule-heavy, document-heavy, spread across several systems and today handled by people or brittle RPA scripts. That is where the first measurable savings will appear, and also where mistakes are expensive. The business case will be won or lost on exception handling, not on the happy path.",{"type":148,"level":374,"text":391},"Software vendors: the agent is the user",{"type":133,"content":393},[394,395,399,400,404],"If a dot does the clicking, the dot is your user. Dots reach apps through OpenAI's plugin ecosystem and otherwise use the browser on their own computer. Products with a clean plugin or API surface will be used well; products that only work through a human-shaped interface will be used badly, or skipped. It is the same shift I described for ",{"tag":284,"to":396,"children":397},"\u002Fblog\u002Fwebmcp-agent-ready-website-guide",[398],"agent-ready websites with WebMCP"," and for ",{"tag":284,"to":401,"children":402},"\u002Fblog\u002Fagentic-commerce-protocols-ucp-acp-guide",[403],"agentic commerce protocols",", now arriving through the largest distribution channel in AI: OpenAI says ChatGPT has 1.2 billion weekly users.",{"type":133,"content":406},[407],"There is a pricing consequence too. Seat-based SaaS assumes one licence per human operator. When one dot does the routine work of several people in a tool, vendors will see fewer seats and heavier usage, and many will move to usage- or outcome-based pricing. The vendors that make their product safe for a dot to operate, with scoped tokens, clear action semantics and approval hooks, will be the ones an IT department allows dots to touch.",{"type":148,"level":374,"text":409},"Developers: more pull requests, the same reviewers",{"type":133,"content":411},[412,413,417,418,288],"OpenAI's headline developer example is a dot that watches customer feedback and opens tested pull requests with video documentation. That is useful, and it widens the gap the industry already has: generating changes is cheap, reviewing them is not. A team that lets dots work on a repository needs a review policy first, with size budgets, required tests and a named human owner for every change. Tasks dots start in Codex or ChatGPT Work count against normal usage limits, so cost control belongs in the same policy. For the engineering side of running agents reliably, see ",{"tag":284,"to":414,"children":415},"\u002Fblog\u002Fharness-engineering-coding-agents",[416],"harness engineering"," and ",{"tag":284,"to":419,"children":420},"\u002Fblog\u002Fagent-loop-explained",[421],"how the agent loop works",{"type":148,"level":374,"text":423},"Europe: in through the company door",{"type":133,"content":425},[426],"The rollout map is unusual. Pro users in the European Economic Area, Switzerland and the UK are excluded at launch, while Business Premium users get dots in every supported region. OpenAI has not said why. Whatever the reason, the effect is clear: in Europe, dots will not spread through employees' personal subscriptions first. They will arrive through the company account, which means through IT, procurement and the data protection officer.",{"type":133,"content":428},[429,430,434,435,288],"That is good news, provided those three are ready. An always-on agent that reads CRM records, inboxes and documents is personal data processing at scale. It needs a legal basis, a data processing agreement, retention rules and, in most cases, a data protection impact assessment. Business, Enterprise and Edu content is not used for training by default, but limited human review can still happen in safety cases, and a dot's memories cannot be inspected one by one, which will make access and erasure requests awkward. Where a dot writes to people on your behalf, the transparency duties of the EU AI Act may also apply. I covered the groundwork in ",{"tag":284,"to":431,"children":432},"\u002Fblog\u002Fgdpr-llm-api-eu-data-residency",[433],"GDPR and EU data residency for LLM APIs"," and in the ",{"tag":284,"to":436,"children":437},"\u002Fblog\u002Feu-ai-act-article-50-developer-checklist",[438],"AI Act Article 50 checklist",{"type":148,"level":374,"text":440},"The market: the personal agent is the new platform war",{"type":133,"content":442},[443],"Dots arrived three weeks after Meta's Muse, which topped the App Store within days, and one day after Instinct, a startup building a personal agent, raised $1 billion at a $10 billion valuation. Meta is going after consumers; OpenAI, at least with this first release, is going after work. The prize is the same: whoever holds the agent that knows your preferences, tools and history holds the relationship, and every other app becomes a supplier to it. That is why memory and integrations, not benchmark scores, will decide this round.",{"type":201,"head":445,"rows":452},[446,448,450],[447],"Who",[449],"What changes first",[451],"What to prepare",[453,460,467,474,481],[454,456,458],[455],"Knowledge workers",[457],"Routine follow-ups move to a dot; the job becomes delegation and review",[459],"Clear briefs, a definition of done, protected review time",[461,463,465],[462],"IT and security",[464],"Agents become identities with credentials and hardware",[466],"Agent IAM, least privilege, offboarding, logs outside the vendor",[468,470,472],[469],"Software vendors",[471],"The agent becomes the operator of the product",[473],"Plugins and APIs, scoped tokens, approval hooks, usage pricing",[475,477,479],[476],"Developers",[478],"More agent-written pull requests",[480],"Review budgets, required tests, a human owner per change",[482,484,486],[483],"European companies",[485],"Access comes through the business account",[487],"DPIA, processing agreement, rules for regulated data",{"type":148,"level":149,"id":123,"text":124},{"type":133,"content":490},[491],"The Enterprise beta is off by default, which gives most organisations a rare moment: the decision can be made before the tool is in use, not after. This is the order I would work in:",{"type":177,"ordered":493,"items":494},true,[495,500,505,510,515,520],[496,499],{"tag":137,"children":497},[498],"Pick one bounded process and measure it."," Invoice follow-ups, proposal updates or support triage. Record today's cycle time and error rate, run a dot on it for four weeks, and measure the same numbers plus the review time it costs.",[501,504],{"tag":137,"children":502},[503],"Write Custom Rules before connecting apps."," Require approval by default for anything that sends, pays, deletes or shares, and loosen a rule only where the activity log shows the dot is reliable.",[506,509],{"tag":137,"children":507},[508],"Treat each dot as an identity."," A separate account, least-privilege scopes, an owner, an expiry date and an offboarding step. Never give a dot a person's credentials.",[511,514],{"tag":137,"children":512},[513],"Keep regulated data out until you control exposure."," Health, payment and HR systems stay disconnected until you know what the model receives, not only what it may do.",[516,519],{"tag":137,"children":517},[518],"Budget review capacity, not just licences."," Every hour a dot saves creates some minutes of checking. Decide who does it and when, or approval fatigue will decide for you.",[521,524],{"tag":137,"children":522},[523],"If you sell software, make it agent-operable."," A plugin or a well-scoped API with clear action semantics is now a distribution channel.",{"type":133,"content":526},[527],"None of this requires betting on OpenAI. Meta, Google and Anthropic are building the same category, and the same controls apply to all of them. The vendor may change; the governance you build now will not.",{"type":148,"level":149,"id":126,"text":127},{"type":133,"content":530},[531],"Dots are the first mass-market product that treats an AI model as a member of staff rather than a feature. The capabilities were already there; OpenAI has packaged them with an identity, a memory, a computer and a price model that looks like labour. That is why the impact will be organisational before it is technical.",{"type":133,"content":533},[534],"The open question is not whether dots can do the work. In a narrow, well-scoped process they clearly can. The question is whether organisations can absorb work that arrives faster than they can verify it, and whether the safeguards around a model whose successor was just held back for leaving its scope will hold once dots reach the wider ChatGPT user base. The companies that come out ahead will treat delegation as a discipline: clear goals, narrow access, real review.",{"type":148,"level":149,"id":129,"text":130},{"type":177,"ordered":493,"items":537},[538,541,544,547,550,553,556,559,562,565,568,571],[539],{"tag":172,"href":36,"children":540},[35],[542],{"tag":172,"href":39,"children":543},[38],[545],{"tag":172,"href":42,"children":546},[41],[548],{"tag":172,"href":45,"children":549},[44],[551],{"tag":172,"href":48,"children":552},[47],[554],{"tag":172,"href":51,"children":555},[50],[557],{"tag":172,"href":54,"children":558},[53],[560],{"tag":172,"href":57,"children":561},[56],[563],{"tag":172,"href":60,"children":564},[59],[566],{"tag":172,"href":63,"children":567},[62],[569],{"tag":172,"href":66,"children":570},[65],[572],{"tag":172,"href":69,"children":573},[68],[575,673,729,780],{"slug":576,"published":577,"minutes":578,"category":7,"tags":579,"keywords":584,"about":596,"sources":606,"cover":667,"og":668,"expertise":72,"locales":669,"lang":74,"title":670,"description":671,"coverAlt":672},"token-saving-tools-coding-agents-top-20","2026-09-28",17,[580,581,582,583],"Claude Code","token usage","context engineering","developer tools",[585,586,587,588,589,590,591,592,593,594,595,582],"reduce Claude Code token usage","token saving tools for coding agents","rtk token killer","lean-ctx","context-mode MCP","Serena MCP","Repomix compress","MCP tool search","prompt caching","ccusage","coding agent cost",[597,600,603],{"name":598,"url":599},"Large language model","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FLarge_language_model",{"name":601,"url":602},"Claude (language model)","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FClaude_(language_model)",{"name":604,"url":605},"Prompt engineering","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FPrompt_engineering",[607,610,613,616,619,622,625,628,631,634,637,640,643,646,649,652,655,658,661,664],{"title":608,"url":609},"rtk: a CLI proxy that filters shell output for coding agents (GitHub)","https:\u002F\u002Fgithub.com\u002Frtk-ai\u002Frtk",{"title":611,"url":612},"lean-ctx: context read modes and shell compression for coding agents (GitHub)","https:\u002F\u002Fgithub.com\u002Fyvgude\u002Flean-ctx",{"title":614,"url":615},"context-mode: sandboxed tool output with SQLite FTS5 search (GitHub)","https:\u002F\u002Fgithub.com\u002Fmksglu\u002Fcontext-mode",{"title":617,"url":618},"Serena: semantic code retrieval and editing over MCP (GitHub)","https:\u002F\u002Fgithub.com\u002Foraios\u002Fserena",{"title":620,"url":621},"token-savior: symbol index, memory and bash compaction over MCP (GitHub)","https:\u002F\u002Fgithub.com\u002FMibayy\u002Ftoken-savior",{"title":623,"url":624},"code-review-graph: a code graph for blast-radius reviews (GitHub)","https:\u002F\u002Fgithub.com\u002Ftirth8205\u002Fcode-review-graph",{"title":626,"url":627},"Aider documentation: repository map","https:\u002F\u002Faider.chat\u002Fdocs\u002Frepomap.html",{"title":629,"url":630},"Repomix: pack a repository into one AI-friendly file (GitHub)","https:\u002F\u002Fgithub.com\u002Fyamadashy\u002Frepomix",{"title":632,"url":633},"Context7: up-to-date library documentation for LLMs (GitHub)","https:\u002F\u002Fgithub.com\u002Fupstash\u002Fcontext7",{"title":635,"url":636},"claude-context: hybrid code search MCP (GitHub)","https:\u002F\u002Fgithub.com\u002Fzilliztech\u002Fclaude-context",{"title":638,"url":639},"caveman: terse output modes for coding agents (GitHub)","https:\u002F\u002Fgithub.com\u002FJuliusBrussee\u002Fcaveman",{"title":641,"url":642},"claude-token-efficient: an eight-rule CLAUDE.md (GitHub)","https:\u002F\u002Fgithub.com\u002Fdrona23\u002Fclaude-token-efficient",{"title":644,"url":645},"claude-code-router: a local model gateway for coding agents (GitHub)","https:\u002F\u002Fgithub.com\u002Fmusistudio\u002Fclaude-code-router",{"title":647,"url":648},"ccusage: token and cost reports from local agent logs (GitHub)","https:\u002F\u002Fgithub.com\u002Fryoppippi\u002Fccusage",{"title":650,"url":651},"LLMLingua: prompt compression (Microsoft, GitHub)","https:\u002F\u002Fgithub.com\u002Fmicrosoft\u002FLLMLingua",{"title":653,"url":654},"ComputingForGeeks: tools that reduce Claude Code token usage, tested (April 2026)","https:\u002F\u002Fcomputingforgeeks.com\u002Freduce-claude-code-token-usage-tools\u002F",{"title":656,"url":657},"Claude Code docs: manage costs effectively","https:\u002F\u002Fcode.claude.com\u002Fdocs\u002Fen\u002Fcosts",{"title":659,"url":660},"Claude Code docs: MCP, output limits and tool search","https:\u002F\u002Fcode.claude.com\u002Fdocs\u002Fen\u002Fmcp",{"title":662,"url":663},"Claude API docs: tool search tool","https:\u002F\u002Fplatform.claude.com\u002Fdocs\u002Fen\u002Fagents-and-tools\u002Ftool-use\u002Ftool-search-tool",{"title":665,"url":666},"Claude API docs: prompt caching","https:\u002F\u002Fplatform.claude.com\u002Fdocs\u002Fen\u002Fbuild-with-claude\u002Fprompt-caching","\u002Fimages\u002Fblog\u002Ftoken-saving-tools-coding-agents-top-20\u002Fcover.webp","\u002Fimages\u002Fblog\u002Ftoken-saving-tools-coding-agents-top-20\u002Fog.jpg",[74,75,76],"Top 20 ways to cut coding-agent tokens: rtk, lean-ctx, Serena and more, ranked by evidence","rtk, lean-ctx, context-mode, Serena and 16 more token savers for coding agents, ranked by evidence, with my own measurements on a real Nuxt codebase.","Bar chart falling from a 15,100-token build log to about 370 tokens after filtering, under the heading Top 20 token savers.",{"slug":674,"published":675,"minutes":676,"category":7,"tags":677,"keywords":682,"about":693,"sources":698,"cover":723,"og":724,"expertise":72,"locales":725,"lang":74,"title":726,"description":727,"coverAlt":728},"agent-loop-explained","2026-09-27",10,[678,679,680,580,681],"Agent loop","Coding agents","Tool calling","Stop conditions",[683,684,685,686,687,688,689,690,691,692],"agent loop","agentic loop","AI agent loop","tool calling loop","Ralph loop coding agent","Claude Code \u002Floop","Claude Code \u002Fgoal","how to stop an AI agent from looping","ReAct agent pattern","stop_reason tool_use",[694,695,696],{"name":31,"url":32},{"name":598,"url":599},{"name":580,"url":697},"https:\u002F\u002Fcode.claude.com\u002Fdocs\u002Fen\u002Foverview",[699,702,705,708,711,714,717,720],{"title":700,"url":701},"Anthropic: Building effective agents (Dec 2024)","https:\u002F\u002Fwww.anthropic.com\u002Fengineering\u002Fbuilding-effective-agents",{"title":703,"url":704},"Yao et al.: ReAct: Synergizing Reasoning and Acting in Language Models (2022)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2210.03629",{"title":706,"url":707},"Claude API docs: Handling stop reasons","https:\u002F\u002Fplatform.claude.com\u002Fdocs\u002Fen\u002Fbuild-with-claude\u002Fhandling-stop-reasons",{"title":709,"url":710},"Anthropic: Effective harnesses for long-running agents (Nov 2025)","https:\u002F\u002Fwww.anthropic.com\u002Fengineering\u002Feffective-harnesses-for-long-running-agents",{"title":712,"url":713},"Simon Willison: Red\u002Fgreen TDD (Agentic Engineering Patterns)","https:\u002F\u002Fsimonwillison.net\u002Fguides\u002Fagentic-engineering-patterns\u002Fred-green-tdd\u002F",{"title":715,"url":716},"Geoffrey Huntley: Ralph Wiggum as a software engineer (Jul 2025)","https:\u002F\u002Fghuntley.com\u002Fralph\u002F",{"title":718,"url":719},"Claude Code docs: Run prompts on a schedule (\u002Floop)","https:\u002F\u002Fcode.claude.com\u002Fdocs\u002Fen\u002Fscheduled-tasks",{"title":721,"url":722},"Claude Code docs: Keep Claude working toward a goal (\u002Fgoal)","https:\u002F\u002Fcode.claude.com\u002Fdocs\u002Fen\u002Fgoal","\u002Fimages\u002Fblog\u002Fagent-loop-explained\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fagent-loop-explained\u002Fog.jpg",[74,75,76],"The agent loop, explained: how coding agents run, and how to make them stop","How the agent loop works in code, which stop conditions and budgets to enforce, and how outer loops like Ralph and Claude Code \u002Floop and \u002Fgoal behave.","A five-step cycle: context, model, tool call, result and a stop check that either ends the agent loop or feeds back into the context.",{"slug":730,"published":675,"minutes":676,"category":7,"tags":731,"keywords":737,"about":747,"sources":755,"cover":774,"og":775,"expertise":72,"locales":776,"lang":74,"title":777,"description":778,"coverAlt":779},"mcp-tool-design-lessons-jira-server",[732,733,734,735,736],"MCP","Tool design","Context engineering","Jira","Agents",[738,739,740,741,742,743,744,745,746],"MCP tool design","MCP best practices","MCP tool descriptions","agent tool selection","MCP context bloat","how many tools should an MCP server have","MCP tool definition token cost","MCP error handling isError","Jira MCP server",[748,751,754],{"name":749,"url":750},"Model Context Protocol","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FModel_Context_Protocol",{"name":752,"url":753},"Jira (software)","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FJira_(software)",{"name":31,"url":32},[756,759,762,765,768,771],{"title":757,"url":758},"Writing effective tools for agents – with agents (Anthropic)","https:\u002F\u002Fwww.anthropic.com\u002Fengineering\u002Fwriting-tools-for-agents",{"title":760,"url":761},"Introducing advanced tool use on the Claude Developer Platform (Anthropic)","https:\u002F\u002Fwww.anthropic.com\u002Fengineering\u002Fadvanced-tool-use",{"title":763,"url":764},"Code execution with MCP (Anthropic)","https:\u002F\u002Fwww.anthropic.com\u002Fengineering\u002Fcode-execution-with-mcp",{"title":766,"url":767},"MCP vs CLI: context window cost (Blocks.ai)","https:\u002F\u002Fblocks.ai\u002Fblog\u002Fmcp-vs-cli-context-window-cost",{"title":769,"url":770},"Demystifying evals for AI agents (Anthropic)","https:\u002F\u002Fwww.anthropic.com\u002Fengineering\u002Fdemystifying-evals-for-ai-agents",{"title":772,"url":773},"MCP 2026-07-28 specification: Tools","https:\u002F\u002Fmodelcontextprotocol.io\u002Fspecification\u002F2026-07-28\u002Fserver\u002Ftools","\u002Fimages\u002Fblog\u002Fmcp-tool-design-lessons-jira-server\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fmcp-tool-design-lessons-jira-server\u002Fog.jpg",[74,75,76],"MCP tool design: lessons from a 20-tool Jira server","MCP tool design that agents get right: token cost of tool definitions, when to merge tools, naming, concise output, errors that steer and a small selection eval.","Network diagram with a Jira MCP server at the hub and five satellites: search, create, transition, comments and test runs",{"slug":781,"published":675,"minutes":782,"category":7,"tags":783,"keywords":788,"about":797,"sources":806,"cover":829,"og":830,"expertise":72,"locales":831,"lang":74,"title":832,"description":833,"coverAlt":834},"harness-engineering-coding-agents",8,[784,679,785,786,787],"Harness engineering","Code quality","Mutation testing","TDD",[416,789,790,791,792,793,794,795,796],"harness engineering coding agents","AI code quality","coding agent guardrails","mutation testing AI generated tests","red\u002Fgreen TDD with AI agents","guides and sensors coding agents","how to make AI agent pull requests mergeable","can I trust tests written by AI",[798,800,803],{"name":786,"url":799},"https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FMutation_testing",{"name":801,"url":802},"Test-driven development","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FTest-driven_development",{"name":804,"url":805},"Static program analysis","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FStatic_program_analysis",[807,810,813,816,819,820,823,826],{"title":808,"url":809},"Birgitta Böckeler: Harness engineering for coding agent users (Apr 2026)","https:\u002F\u002Fmartinfowler.com\u002Farticles\u002Fharness-engineering.html",{"title":811,"url":812},"Birgitta Böckeler: Maintainability sensors for coding agents (May 2026)","https:\u002F\u002Fmartinfowler.com\u002Farticles\u002Fsensors-for-coding-agents.html",{"title":814,"url":815},"Simon Willison: Agentic Engineering Patterns","https:\u002F\u002Fsimonwillison.net\u002Fguides\u002Fagentic-engineering-patterns\u002F",{"title":817,"url":818},"Simon Willison: First run the tests","https:\u002F\u002Fsimonwillison.net\u002Fguides\u002Fagentic-engineering-patterns\u002Ffirst-run-the-tests\u002F",{"title":709,"url":710},{"title":821,"url":822},"Claude Code docs: How Claude remembers your project","https:\u002F\u002Fcode.claude.com\u002Fdocs\u002Fen\u002Fmemory",{"title":824,"url":825},"Stryker Mutator documentation","https:\u002F\u002Fstryker-mutator.io\u002Fdocs\u002F",{"title":827,"url":828},"Infection: command line options","https:\u002F\u002Finfection.github.io\u002Fguide\u002Fcommand-line-options.html","\u002Fimages\u002Fblog\u002Fharness-engineering-coding-agents\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fharness-engineering-coding-agents\u002Fog.jpg",[74,75,76],"Harness engineering: guides and sensors that make agent PRs mergeable","Harness engineering for coding agents: guides and sensors, where to run each check, red\u002Fgreen TDD, and mutation testing to verify the tests the agent wrote.","Concentric rings around a coding agent's model: behaviour, architecture fitness and maintainability harnesses, from outside in.",1790970925191]