[{"data":1,"prerenderedAt":881},["ShallowReactive",2],{"blog-ai-agent-memory-design-en":3},{"slug":4,"published":5,"minutes":6,"category":7,"tags":8,"keywords":13,"about":24,"sources":34,"cover":83,"og":84,"expertise":85,"locales":86,"lang":87,"title":90,"description":91,"coverAlt":92,"metaTitle":93,"takeaways":94,"faq":100,"toc":119,"blocks":153,"others":600},"ai-agent-memory-design","2026-10-02",13,"agents",[9,10,11,12],"AI agent memory","Context engineering","Memory poisoning","GDPR",[14,15,16,17,18,19,20,21,22,23],"AI agent memory design","long-term memory for AI agents","episodic semantic procedural memory LLM","agent memory architecture","ChatGPT memory vs Claude memory","Claude memory tool","AI memory poisoning","LLM context compaction","AI agent memory GDPR","short-term vs long-term memory agents",[25,28,31],{"name":26,"url":27},"Intelligent agent","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FIntelligent_agent",{"name":29,"url":30},"General Data Protection Regulation","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FGeneral_Data_Protection_Regulation",{"name":32,"url":33},"Prompt injection","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FPrompt_injection",[35,38,41,44,47,50,53,56,59,62,65,68,71,74,77,80],{"title":36,"url":37},"Anthropic docs: Memory tool","https:\u002F\u002Fplatform.claude.com\u002Fdocs\u002Fen\u002Fagents-and-tools\u002Ftool-use\u002Fmemory-tool",{"title":39,"url":40},"Anthropic docs: Context editing","https:\u002F\u002Fplatform.claude.com\u002Fdocs\u002Fen\u002Fbuild-with-claude\u002Fcontext-editing",{"title":42,"url":43},"Anthropic Engineering: Effective context engineering for AI agents","https:\u002F\u002Fwww.anthropic.com\u002Fengineering\u002Feffective-context-engineering-for-ai-agents",{"title":45,"url":46},"Sumers et al.: Cognitive Architectures for Language Agents (CoALA)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2309.02427",{"title":48,"url":49},"Packer et al.: MemGPT, Towards LLMs as Operating Systems","https:\u002F\u002Farxiv.org\u002Fabs\u002F2310.08560",{"title":51,"url":52},"Park et al.: Generative Agents, Interactive Simulacra of Human Behavior","https:\u002F\u002Farxiv.org\u002Fabs\u002F2304.03442",{"title":54,"url":55},"Unit 42: When AI Remembers Too Much, persistent behaviors in agents memory","https:\u002F\u002Funit42.paloaltonetworks.com\u002Findirect-prompt-injection-poisons-ai-longterm-memory\u002F",{"title":57,"url":58},"From Untrusted Input to Trusted Memory: A Systematic Study of Memory Poisoning Attacks in LLM Agents (preprint)","https:\u002F\u002Farxiv.org\u002Fhtml\u002F2606.04329v1",{"title":60,"url":61},"The Hacker News: ChatGPT macOS flaw could have enabled long-term spyware via memory function","https:\u002F\u002Fthehackernews.com\u002F2024\u002F09\u002Fchatgpt-macos-flaw-couldve-enabled-long.html",{"title":63,"url":64},"Vectorize: OWASP ASI06, Memory and Context Poisoning explained","https:\u002F\u002Fvectorize.io\u002Farticles\u002Fowasp-asi06",{"title":66,"url":67},"Claude Help Center: Use chat search and memory to build on previous context","https:\u002F\u002Fsupport.claude.com\u002Fen\u002Farticles\u002F11817273-use-claude-s-chat-search-and-memory-to-build-on-previous-context",{"title":69,"url":70},"OpenAI Help Center: Memory in ChatGPT","https:\u002F\u002Fhelp.openai.com\u002Fen\u002Farticles\u002F8590148-memory-faq",{"title":72,"url":73},"OpenAI Help Center: Dots privacy, security, and safety FAQs","https:\u002F\u002Fhelp.openai.com\u002Fen\u002Farticles\u002F20001529-dots-privacy-security-and-safety-faqs",{"title":75,"url":76},"Flavio Copes: A deep dive into OpenAI dots (quotes the dots documentation on memory)","https:\u002F\u002Fflaviocopes.com\u002Fopenai-dots\u002F",{"title":78,"url":79},"GDPR Article 5: Principles relating to processing of personal data","https:\u002F\u002Fgdpr-info.eu\u002Fart-5-gdpr\u002F",{"title":81,"url":82},"GDPR Article 17: Right to erasure","https:\u002F\u002Fgdpr-info.eu\u002Fart-17-gdpr\u002F","\u002Fimages\u002Fblog\u002Fai-agent-memory-design\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fai-agent-memory-design\u002Fog.jpg","ai-engineer",[87,88,89],"en","de","hu","Designing memory for AI agents: tiers, write rules, poisoning and GDPR","How to design AI agent memory: context vs session vs long-term tiers, what to write and never store, retrieval, compaction, poisoning and GDPR erasure.","Diagram: nested memory layers of an AI agent, from the working context window through session state to long-term episodic and semantic memory.","Memory design for AI agents: tiers and risks · Balázs Csorba",[95,96,97,98,99],"Treat agent memory as three tiers with different lifetimes: working context (the window), session state (one task or thread) and long-term memory (across sessions). Each tier needs its own write and delete rules.","The hard part is the write path, not the store. Decide what is worth remembering, from which sources, with provenance, a scope and an expiry, and never let untrusted content write memory unchecked.","Memory is a persistence layer for prompt injection: one poisoned write can steer every later session, and summarisation and compaction are write channels too.","For GDPR, every memory must be attributable to a person and deletable together with its derived copies (summaries, embeddings, caches). A store you can only reset as a whole is a compliance problem.","Products differ sharply: ChatGPT and Claude let users view, edit and delete memories, the Claude memory tool leaves storage to you, and OpenAI dots let you delete only the whole dot.",[101,104,107,110,113,116],{"q":102,"a":103},"What is memory in an AI agent?","It is anything an agent can use in a later step or session that is not part of the model weights: the current context window, state kept for the running task, and a persistent store such as files or a database that the agent reads and writes across sessions. LLMs are stateless, so every form of memory is something your system writes into the prompt.",{"q":105,"a":106},"What is the difference between short-term and long-term memory for AI agents?","Short-term memory is the working context of the current conversation or task and disappears with it, or is compacted. Long-term memory is stored outside the model, survives across sessions and is retrieved on demand. In between sits session state, such as a progress file or thread summary, that lives for the duration of one piece of work.",{"q":108,"a":109},"What are episodic, semantic and procedural memory in LLM agents?","The CoALA framework borrows these terms from cognitive science. Episodic memory stores specific experiences (what happened in a past task), semantic memory stores general facts (the customer prefers email), and procedural memory stores learned skills or ways of acting. Working memory is the active context on top of them.",{"q":111,"a":112},"What should an AI agent never store in memory?","Secrets and credentials, government IDs and financial account numbers, special-category personal data unless you have a clear basis and purpose, raw tool output and web content that could carry instructions, and anything the user did not expect to be kept. Claude, for example, excludes government IDs, financial account numbers, criminal history and immigration status by default.",{"q":114,"a":115},"What is AI memory poisoning?","It is an attack where adversarial content gets written into an agent's persistent memory, so that it influences behaviour in later sessions. OWASP lists it as ASI06, Memory and Context Poisoning, in its Top 10 for Agentic Applications. Unlike ordinary prompt injection, it does not end when the session does.",{"q":117,"a":118},"Is AI agent memory compatible with GDPR?","It can be, if you design for it. Memory that holds personal data must follow purpose limitation, data minimisation, accuracy and storage limitation (Article 5), and you need to find and erase a person's memories on request (Article 17). That requires per-user scoping, provenance and deletion that also reaches summaries and embeddings.",[120,123,126,129,132,135,138,141,144,147,150],{"id":121,"title":122},"memory-tiers","Three tiers: working context, session, long-term",{"id":124,"title":125},"memory-types","Episodic, semantic, procedural: what you actually store",{"id":127,"title":128},"write-policy","What to write, when, and what never to store",{"id":130,"title":131},"retrieval","Retrieving memories without drowning the context",{"id":133,"title":134},"compaction","Compaction and summarisation: lossy by design",{"id":136,"title":137},"memory-poisoning","Memory poisoning: the injection that stays",{"id":139,"title":140},"gdpr","GDPR: access, erasure and the derived-data problem",{"id":142,"title":143},"products","How current products handle memory",{"id":145,"title":146},"checklist","A memory design checklist",{"id":148,"title":149},"recommendation","My recommendation",{"id":151,"title":152},"sources","Sources",[154,158,172,175,178,181,190,193,194,197,251,259,260,263,270,273,276,305,306,309,340,348,349,352,355,358,359,362,365,372,378,379,382,414,417,418,421,512,519,520,523,544,545,548,549],{"type":155,"content":156},"paragraph",[157],"Large language models are stateless. Everything an agent \"remembers\" is text that your system decided to put back into the prompt, and the interesting engineering questions are therefore about the loop around the model: what gets written, where, by whom, how it is found again, and how it is removed. Memory is also where agents stop being a demo. A coding agent that re-learns your repository every morning, or a support agent that asks the same question for the third time, is a memory problem.",{"type":155,"content":159},[160,161,166,167,171],"Memory is also an attack surface and a data-protection liability, which most tutorials skip. In this article I walk through the tiers I use, the three memory types, a write policy with an explicit \"never store\" list, retrieval, compaction, poisoning and GDPR, and then compare how ChatGPT, Claude and OpenAI dots handle it today. It builds on ",{"tag":162,"to":163,"children":164},"link","\u002Fblog\u002Fagent-loop-explained",[165],"how the agent loop works"," and on the threat model in ",{"tag":162,"to":168,"children":169},"\u002Fblog\u002Fprompt-injection-lethal-trifecta-patterns",[170],"the lethal trifecta article",".",{"type":173,"level":174,"id":121,"text":122},"heading",2,{"type":155,"content":176},[177],"I find it useful to separate memory by lifetime, because lifetime decides who may write it, how it is deleted and what it costs. The research literature agrees on the shape. MemGPT framed the problem as an operating system: it \"intelligently manages different memory tiers\" to give the model an extended context, moving information between fast and slow storage like RAM and disk.",{"type":155,"content":179},[180],"Working context is the window itself. It is expensive and it degrades: Anthropic describes context rot, where \"as the number of tokens in the context window increases, the model's ability to accurately recall information from that context decreases\". Session state is what you keep for one task or thread so that an interruption does not destroy progress, for example a progress file or a conversation summary. Long-term memory survives across sessions and is the only tier where privacy, poisoning and erasure become permanent problems.",{"type":182,"attrs":183,"inner":187,"caption":188},"diagram",{"viewBox":184,"role":185,"aria-labelledby":186},"0 0 720 306","img","d1-mem-t d1-mem-d","\u003Ctitle id=\"d1-mem-t\">Memory tiers of an agent\u003C\u002Ftitle>\u003Cdesc id=\"d1-mem-d\">Three boxes from left to right: working context, session state and long-term memory. An upper arrow shows writes through a gate from left to right, a lower arrow shows retrieval from right to left. Below, long-term memory splits into episodic, semantic and procedural memory.\u003C\u002Fdesc>\u003Ctext x=\"20\" y=\"28\" class=\"d-title\">Memory tiers of an agent\u003C\u002Ftext>\u003Ctext x=\"700\" y=\"28\" text-anchor=\"end\" class=\"d-label\">lifetime grows to the right\u003C\u002Ftext>\u003Crect x=\"20\" y=\"62\" width=\"200\" height=\"84\" rx=\"10\" class=\"d-sky\" \u002F>\u003Ctext x=\"120\" y=\"100\" text-anchor=\"middle\" class=\"d-text\">Working context\u003C\u002Ftext>\u003Ctext x=\"120\" y=\"121\" text-anchor=\"middle\" class=\"d-small\">the context window\u003C\u002Ftext>\u003Crect x=\"260\" y=\"62\" width=\"200\" height=\"84\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"360\" y=\"100\" text-anchor=\"middle\" class=\"d-text\">Session state\u003C\u002Ftext>\u003Ctext x=\"360\" y=\"121\" text-anchor=\"middle\" class=\"d-small\">one task or thread\u003C\u002Ftext>\u003Crect x=\"500\" y=\"62\" width=\"200\" height=\"84\" rx=\"10\" class=\"d-accent\" \u002F>\u003Ctext x=\"600\" y=\"100\" text-anchor=\"middle\" class=\"d-text\">Long-term memory\u003C\u002Ftext>\u003Ctext x=\"600\" y=\"121\" text-anchor=\"middle\" class=\"d-small\">across sessions\u003C\u002Ftext>\u003Cpath d=\"M220 92 H252\" class=\"d-line\" \u002F>\u003Cpath d=\"M260 92 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Cpath d=\"M460 92 H492\" class=\"d-line\" \u002F>\u003Cpath d=\"M500 92 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Cpath d=\"M260 118 H228\" class=\"d-line\" \u002F>\u003Cpath d=\"M220 118 l9 -5 v10 z\" class=\"d-head\" \u002F>\u003Cpath d=\"M500 118 H468\" class=\"d-line\" \u002F>\u003Cpath d=\"M460 118 l9 -5 v10 z\" class=\"d-head\" \u002F>\u003Ctext x=\"20\" y=\"176\" class=\"d-label\">→ write, through a gate     ← retrieve on demand\u003C\u002Ftext>\u003Ctext x=\"20\" y=\"214\" class=\"d-title\">Long-term memory splits by kind\u003C\u002Ftext>\u003Crect x=\"20\" y=\"226\" width=\"215\" height=\"58\" rx=\"10\" class=\"d-sky\" \u002F>\u003Ctext x=\"127.5\" y=\"251\" text-anchor=\"middle\" class=\"d-text\">Episodic\u003C\u002Ftext>\u003Ctext x=\"127.5\" y=\"272\" text-anchor=\"middle\" class=\"d-small\">what happened\u003C\u002Ftext>\u003Crect x=\"252\" y=\"226\" width=\"215\" height=\"58\" rx=\"10\" class=\"d-gold\" \u002F>\u003Ctext x=\"359.5\" y=\"251\" text-anchor=\"middle\" class=\"d-text\">Semantic\u003C\u002Ftext>\u003Ctext x=\"359.5\" y=\"272\" text-anchor=\"middle\" class=\"d-small\">what is true\u003C\u002Ftext>\u003Crect x=\"485\" y=\"226\" width=\"215\" height=\"58\" rx=\"10\" class=\"d-mint\" \u002F>\u003Ctext x=\"592.5\" y=\"251\" text-anchor=\"middle\" class=\"d-text\">Procedural\u003C\u002Ftext>\u003Ctext x=\"592.5\" y=\"272\" text-anchor=\"middle\" class=\"d-small\">how to act\u003C\u002Ftext>",[189],"Lifetime grows from left to right, and so does the cost of a bad write. Only the gate between session and long-term memory decides what becomes permanent.",{"type":155,"content":191},[192],"The consequence for design is simple. Working context is rebuilt on every call and can be as noisy as the task needs. Session state should be small, structured and owned by the harness. Long-term memory needs a write gate, because everything that crosses it becomes something the agent will later treat as its own knowledge.",{"type":173,"level":174,"id":124,"text":125},{"type":155,"content":195},[196],"The CoALA paper by Sumers, Yao, Narasimhan and Griffiths organises language agents around working memory plus three long-term types borrowed from cognitive science: episodic, semantic and procedural. The split is useful in practice because each type has a different write trigger, a different retrieval pattern and a different failure mode.",{"type":198,"head":199,"rows":210},"table",[200,202,204,206,208],[201],"Type",[203],"Holds",[205],"Example",[207],"Typical form",[209],"Main failure mode",[211,225,238],[212,217,219,221,223],[213],{"tag":214,"children":215},"strong",[216],"Episodic",[218],"Specific past experiences",[220],"Last Tuesday the refund flow failed because the order was already shipped",[222],"Time-stamped event log or short summaries",[224],"Noise grows without bound; old episodes mislead",[226,230,232,234,236],[227],{"tag":214,"children":228},[229],"Semantic",[231],"General facts and preferences",[233],"Acme Corp prefers email follow-ups",[235],"Key-value facts, profile files, notes",[237],"Stale or wrong facts; contradictions after updates",[239,243,245,247,249],[240],{"tag":214,"children":241},[242],"Procedural",[244],"Learned ways of acting",[246],"For this repo, run the linter before the tests",[248],"Rules, playbooks, prompt snippets, skills",[250],"A poisoned or outdated procedure is executed, not just recalled",{"type":155,"content":252},[253,254,258],"The Generative Agents paper from Stanford and Google shows why episodic memory alone is not enough: the agents store experiences in natural language and \"synthesize those memories over time into higher-level reflections\", which is how events turn into semantic knowledge. I would copy that idea but keep the reflection step inspectable. Procedural memory deserves the most suspicion, since a memory that changes how the agent acts is closer to code than to data. If you let agents write their own skills, treat them like a pull request (see ",{"tag":162,"to":255,"children":256},"\u002Fblog\u002Fcoding-agent-skills-workflow",[257],"coding agent skills",").",{"type":173,"level":174,"id":127,"text":128},{"type":155,"content":261},[262],"Most memory systems fail on the write path. Either the agent writes everything (and retrieval drowns in noise) or it writes on a whim (and important facts are missing). I would define a write policy explicitly instead of leaving it to the model alone. The Claude memory tool documentation points the same way: you can guide what Claude writes, for example by telling it to write down only information relevant to a given topic, and you can add validation that strips sensitive data before your handler persists a file.",{"type":182,"attrs":264,"inner":267,"caption":268},{"viewBox":265,"role":185,"aria-labelledby":266},"0 0 720 250","d2-mem-t d2-mem-d","\u003Ctitle id=\"d2-mem-t\">A write path with gates\u003C\u002Ftitle>\u003Cdesc id=\"d2-mem-d\">Five steps from left to right: candidate memory, source check, minimise and classify, deduplicate and merge, store with metadata. A dashed box below says rejected candidates are dropped or confirmed with the user.\u003C\u002Fdesc>\u003Ctext x=\"20\" y=\"28\" class=\"d-title\">A write path with gates\u003C\u002Ftext>\u003Ctext x=\"700\" y=\"28\" text-anchor=\"end\" class=\"d-label\">every write passes the same gates\u003C\u002Ftext>\u003Crect x=\"14\" y=\"62\" width=\"124\" height=\"76\" rx=\"10\" class=\"d-sky\" \u002F>\u003Ctext x=\"76\" y=\"96\" text-anchor=\"middle\" class=\"d-text\">Candidate\u003C\u002Ftext>\u003Ctext x=\"76\" y=\"117\" text-anchor=\"middle\" class=\"d-small\">user or model\u003C\u002Ftext>\u003Cpath d=\"M138 100 H151\" class=\"d-line\" \u002F>\u003Cpath d=\"M159 100 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Crect x=\"159\" y=\"62\" width=\"124\" height=\"76\" rx=\"10\" class=\"d-accent\" \u002F>\u003Ctext x=\"221\" y=\"96\" text-anchor=\"middle\" class=\"d-text\">Source check\u003C\u002Ftext>\u003Ctext x=\"221\" y=\"117\" text-anchor=\"middle\" class=\"d-small\">who said it?\u003C\u002Ftext>\u003Cpath d=\"M283 100 H296\" class=\"d-line\" \u002F>\u003Cpath d=\"M304 100 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Crect x=\"304\" y=\"62\" width=\"124\" height=\"76\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"366\" y=\"96\" text-anchor=\"middle\" class=\"d-text\">Minimise\u003C\u002Ftext>\u003Ctext x=\"366\" y=\"117\" text-anchor=\"middle\" class=\"d-small\">strip PII and secrets\u003C\u002Ftext>\u003Cpath d=\"M428 100 H441\" class=\"d-line\" \u002F>\u003Cpath d=\"M449 100 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Crect x=\"449\" y=\"62\" width=\"124\" height=\"76\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"511\" y=\"96\" text-anchor=\"middle\" class=\"d-text\">Merge\u003C\u002Ftext>\u003Ctext x=\"511\" y=\"117\" text-anchor=\"middle\" class=\"d-small\">dedupe, resolve\u003C\u002Ftext>\u003Cpath d=\"M573 100 H586\" class=\"d-line\" \u002F>\u003Cpath d=\"M594 100 l-9 -5 v10 z\" class=\"d-head\" \u002F>\u003Crect x=\"594\" y=\"62\" width=\"124\" height=\"76\" rx=\"10\" class=\"d-mint\" \u002F>\u003Ctext x=\"656\" y=\"96\" text-anchor=\"middle\" class=\"d-text\">Store\u003C\u002Ftext>\u003Ctext x=\"656\" y=\"117\" text-anchor=\"middle\" class=\"d-small\">owner, source, TTL\u003C\u002Ftext>\u003Cpath d=\"M221 138 V176\" class=\"d-line d-dash\" \u002F>\u003Cpath d=\"M221 184 l-5 -9 h10 z\" class=\"d-head\" \u002F>\u003Crect x=\"14\" y=\"184\" width=\"692\" height=\"46\" rx=\"10\" class=\"d-box d-dash\" \u002F>\u003Ctext x=\"360\" y=\"212\" text-anchor=\"middle\" class=\"d-small\">Rejected: drop it, or ask the user to confirm\u003C\u002Ftext>",[269],"The source check is the security gate: content from the web, a document or a tool result is evidence, never an instruction to remember.",{"type":155,"content":271},[272],"Write at a few well-defined moments rather than continuously: when the user states a stable preference or correction, at the end of a task (outcome, decisions, open questions), and just before context is cleared or compacted. Anthropic's context editing documentation describes the last one: it pairs with the memory tool so that Claude can save important information to memory before content is cleared. Each entry should carry an owner, a source, a timestamp, a scope (user, project, tenant) and an expiry or review date.",{"type":155,"content":274},[275],"What I would never store, whatever the agent proposes:",{"type":277,"ordered":278,"items":279},"list",false,[280,285,290,295,300],[281,284],{"tag":214,"children":282},[283],"Secrets and credentials."," API keys, tokens and passwords belong in a vault. Anthropic notes that Claude usually refuses to write sensitive information to memory files but recommends your own validation for stronger guarantees. dots go further and say their context does not keep credentials, images or screenshots.",[286,289],{"tag":214,"children":287},[288],"Identifiers and special categories"," such as government IDs, financial account numbers, criminal history and immigration status. Claude excludes exactly these by default, and health, religion, politics and identity topics are off unless the user opts in.",[291,294],{"tag":214,"children":292},[293],"Raw tool output and web content."," Summarise facts you verified, but do not store pages or documents verbatim: they can carry instructions that would then sit in your agent's trusted memory.",[296,299],{"tag":214,"children":297},[298],"Inferences about people the user did not volunteer",", and anything the user would not expect to be kept. If surprise is likely, ask first.",[301,304],{"tag":214,"children":302},[303],"Third-party personal data"," (customers, colleagues) unless you have a defined purpose and legal basis for it.",{"type":173,"level":174,"id":130,"text":131},{"type":155,"content":307},[308],"Memory that cannot be found is just storage. I would start with the simplest pattern that works and only add machinery when evals show you need it. Anthropic calls the underlying idea just-in-time retrieval: instead of loading everything up front, the agent keeps lightweight identifiers such as file paths and loads data when needed. The memory tool follows it, since Claude views the memory directory first and then opens only the relevant files.",{"type":277,"ordered":278,"items":310},[311,316,321,330,335],[312,315],{"tag":214,"children":313},[314],"Scope before search."," Filter by user, tenant and project first, then rank. A cross-tenant memory leak is a data breach, and it is far easier to prevent with a hard filter than with a prompt.",[317,320],{"tag":214,"children":318},[319],"Start with a small index."," One short overview file or profile that is always loaded, plus topic files or records that are fetched on demand, works better than a vector database for a few hundred memories.",[322,325,326,171],{"tag":214,"children":323},[324],"Add hybrid search when volume grows."," Combine keyword and embedding search, rerank, and weight recency; the mechanics are the same as in ",{"tag":162,"to":327,"children":328},"\u002Fblog\u002Frag-pipeline-chunking-hybrid-search-reranking",[329],"a RAG pipeline",[331,334],{"tag":214,"children":332},[333],"Cap the budget."," Inject only a fixed number of memories or tokens, and show the agent where each one came from and when it was written.",[336,339],{"tag":214,"children":337},[338],"Mark memory as data."," Present retrieved memories in a clearly delimited block as context to weigh, not as instructions to obey. This helps but does not remove the poisoning risk below.",{"type":155,"content":341},[342,343,347],"Measure retrieval like any other feature: build a small set of \"would the right memory have been found?\" cases and track hit rate and wrong-memory rate over time (see ",{"tag":162,"to":344,"children":345},"\u002Fblog\u002Fllm-evals-for-product-features",[346],"LLM evals for product features","). Staleness is a retrieval problem too. When a fact changes, update or supersede the old entry; do not append a contradiction and hope the model picks the newer one.",{"type":173,"level":174,"id":133,"text":134},{"type":155,"content":350},[351],"Compaction summarises a conversation that approaches its limit and restarts with the summary. Anthropic describes it as distilling the context window \"in a high-fidelity manner\" and names the central trade-off: overly aggressive compaction risks losing subtle but critical details. It also describes structured note-taking, where the agent regularly writes notes persisted outside the context window and reads them back later. Their Pokémon example kept precise tallies across thousands of game steps this way.",{"type":155,"content":353},[354],"The Claude documentation states the division of labour well: context editing clears specific tool results, compaction summarises the whole conversation on the server, and for long-running agents memory \"preserves the information that must survive summarization\". The default for tool-result clearing triggers at 100,000 input tokens and keeps the last three tool uses, so decisions that matter should be written to memory before that moment, not reconstructed from a summary afterwards.",{"type":155,"content":356},[357],"There is a security consequence. A summary is generated by a model that has just read untrusted content, and its output is then stored and trusted. Compaction is a write channel and must pass the same gates as any other memory write.",{"type":173,"level":174,"id":136,"text":137},{"type":155,"content":360},[361],"Prompt injection normally dies with the session. With memory it does not. In 2024 the researcher Johann Rehberger showed that a malicious document could make ChatGPT store hidden instructions in its long-term memory, after which conversations in new threads kept being sent to an attacker's server. In October 2025 Palo Alto Networks Unit 42 described the same class against Amazon Bedrock Agents: a crafted web page manipulated the session summarisation step, the injected instructions were stored, and later sessions silently exfiltrated user data. Their key observation is that memory contents are injected into the system instructions of orchestration prompts, often prioritised over user input.",{"type":155,"content":363},[364],"A 2026 preprint on memory poisoning systematises this into four write channels: explicit instruction-executed writes, system-prompt-driven writes, compaction-driven writes and experience-to-procedure writes. On the two agents it tested with GPT-OSS-120B, average attack success was 66.67 percent and 34.25 percent, and the four prompt-injection defences it evaluated left significant gaps. OWASP now tracks the class as ASI06, Memory and Context Poisoning. I would not read the exact percentages as a forecast for your system, but the structure of the problem is clear: the more aggressively an agent reads and writes memory, the larger the surface.",{"type":366,"variant":367,"title":368,"body":369},"callout","warn","Design rule",[370],[371],"Content that came from outside the user (web pages, emails, documents, tool results) may be remembered only as a short fact with a source, never as an instruction, preference or procedure. If a memory would change what the agent does rather than what it knows, require a human or a separate check.",{"type":155,"content":373},[374,375,377],"Concretely, I would apply four controls: gate writes by source (the user's own messages are trusted differently from fetched content), give every memory provenance so it can be audited and bulk-removed, review procedural memory like code, and log all reads and writes so that you can answer \"why did the agent do that?\" afterwards. Isolating the agent from outbound channels limits the damage when a bad memory slips through; the patterns in ",{"tag":162,"to":168,"children":376},[170]," apply directly.",{"type":173,"level":174,"id":139,"text":140},{"type":155,"content":380},[381],"As soon as a long-term memory holds personal data, the GDPR principles in Article 5 apply to it. Data must be collected for specified purposes and not processed incompatibly (purpose limitation), be adequate, relevant and limited to what is necessary (data minimisation), be accurate and kept up to date, with inaccurate data erased or rectified without delay, and be kept identifiable no longer than necessary (storage limitation). Article 17 gives data subjects the right to erasure, among other grounds where the data is no longer necessary for its purpose, consent is withdrawn or processing was unlawful. A user asking \"what do you remember about me, and please forget it\" is a normal request you must be able to serve.",{"type":277,"ordered":278,"items":383},[384,389,394,399,404],[385,388],{"tag":214,"children":386},[387],"Attribute every memory to a person."," A user or subject ID on each record makes access and erasure a query instead of an investigation. Memories written about third parties need an identifier too.",[390,393],{"tag":214,"children":391},[392],"Delete derived copies."," Summaries, embeddings, search indexes, caches and backups are copies. Store the source memory ID in each derivative so deleting one cascades to all.",[395,398],{"tag":214,"children":396},[397],"Expire by default."," Storage limitation means a review date or TTL on every entry, and the memory tool documentation lists expiration of long-unused files as a recommended safeguard.",[400,403],{"tag":214,"children":401},[402],"Show and let users correct."," Accuracy is a principle, not a feature. A visible memory view with edit and delete is the cheapest way to meet it.",[405,408,409,413],{"tag":214,"children":406},[407],"Keep logs free of memory content"," and keep the processing location in mind; see ",{"tag":162,"to":410,"children":411},"\u002Fblog\u002Fgdpr-llm-api-eu-data-residency",[412],"GDPR and LLM APIs"," for the data-residency side.",{"type":155,"content":415},[416],"This is where product design differs most, as the next section shows. I would not call any of this legal advice; check your lawful basis and retention periods with your data protection officer. But the engineering requirement is unambiguous: a memory you cannot inspect, correct or delete per item is hard to defend.",{"type":173,"level":174,"id":142,"text":143},{"type":155,"content":419},[420],"The facts below come from the vendors' documentation as of 2 October 2026 or from sources quoting it; the OpenAI help pages blocked my automated access, so for ChatGPT and dots I relied on search extracts of the help articles and on a write-up that quotes the dots documentation. Check the current wording before you rely on a detail.",{"type":198,"head":422,"rows":433},[423,425,427,429,431],[424],"Aspect",[426],"ChatGPT memory",[428],"Claude memory (claude.ai)",[430],"Claude memory tool (API)",[432],"OpenAI dots",[434,447,460,473,486,499],[435,439,441,443,445],[436],{"tag":214,"children":437},[438],"What it is",[440],"Saved memories plus reference to chat history",[442],"Short topics saved from chats, per project",[444],"File operations Claude requests, executed by your app",[446],"Dot-specific notes plus relevant ChatGPT memory",[448,452,454,456,458],[449],{"tag":214,"children":450},[451],"Where it lives",[453],"OpenAI",[455],"Anthropic",[457],"Your infrastructure under \u002Fmemories",[459],"OpenAI, notes separate from ChatGPT memory",[461,465,467,469,471],[462],{"tag":214,"children":463},[464],"View and edit",[466],"Settings, Personalization, Manage memories",[468],"Settings, Memory: view, edit, delete",[470],"Whatever you build",[472],"Cannot view or correct individual memories",[474,478,480,482,484],[475],{"tag":214,"children":476},[477],"Delete",[479],"Per item or all; deleting a chat does not delete its memory",[481],"Per item, pause or reset all",[483],"Your handler decides (delete command, expiry)",[485],"Only by deleting the dot",[487,491,493,495,497],[488],{"tag":214,"children":489},[490],"Sensitive data",[492],"Not covered in the pages I could read",[494],"Excludes IDs, account numbers, criminal and immigration data; health, religion and politics opt-in",[496],"You validate; Claude usually refuses",[498],"Context keeps no credentials, images or screenshots",[500,504,506,508,510],[501],{"tag":214,"children":502},[503],"Off switch",[505],"Yes",[507],"Pause, incognito chats",[509],"You do not enable the tool",[511],"Settings do not necessarily change existing notes",{"type":155,"content":513},[514,515,171],"Three details stand out. ChatGPT keeps a log of deleted saved memories for up to 30 days for safety and debugging, and a memory survives the deletion of the chat it came from. Claude scopes memory per project, enables it by default on Free, Pro and Max plans and leaves it to owners on Team and Enterprise. For dots, the documentation says you cannot view, correct or delete individual dot memories, disconnecting a plugin does not delete what the dot learned from it, and anything the dot added to ChatGPT memory stays after the dot is deleted. For a personal assistant that may be acceptable; for an employer connecting it to customer data it is a GDPR question first. I discuss the wider picture in ",{"tag":162,"to":516,"children":517},"\u002Fblog\u002Fopenai-dots-always-on-agents-impact",[518],"the dots impact analysis",{"type":173,"level":174,"id":145,"text":146},{"type":155,"content":521},[522],"This is the order in which I would work through a new agent:",{"type":277,"ordered":524,"items":525},true,[526,528,530,532,534,536,538,540,542],[527],"Write down which tiers you need. Many agents need only working context plus a session progress file.",[529],"Pick the memory types and give each a different record shape and expiry.",[531],"Define the write policy: triggers, allowed sources, and the \"never store\" list. Put source checks in code, not only in the prompt.",[533],"Add owner, source, timestamp, scope and expiry to every record.",[535],"Scope retrieval by tenant and user before ranking; cap injected tokens.",[537],"Route compaction and summarisation output through the same write gates.",[539],"Review procedural memory like code; require approval for memories that change behaviour.",[541],"Give users a view with edit, delete and pause; make erasure cascade to derived data.",[543],"Log reads and writes, and test with poisoning cases and retrieval evals before launch.",{"type":173,"level":174,"id":148,"text":149},{"type":155,"content":546},[547],"Start small: a session progress file and a short, user-visible profile memory. Add long-term episodic memory only when you can show that it improves outcomes in your evals, and add procedural memory last and with review. The products show where the market is heading, towards agents that remember by default, but they also show the open questions: control, erasure and trust in what was remembered. If you build on the API, the memory tool is a good reference design precisely because it puts storage, validation and deletion in your hands.",{"type":173,"level":174,"id":151,"text":152},{"type":277,"ordered":524,"items":550},[551,555,558,561,564,567,570,573,576,579,582,585,588,591,594,597],[552],{"tag":553,"href":37,"children":554},"a",[36],[556],{"tag":553,"href":40,"children":557},[39],[559],{"tag":553,"href":43,"children":560},[42],[562],{"tag":553,"href":46,"children":563},[45],[565],{"tag":553,"href":49,"children":566},[48],[568],{"tag":553,"href":52,"children":569},[51],[571],{"tag":553,"href":55,"children":572},[54],[574],{"tag":553,"href":58,"children":575},[57],[577],{"tag":553,"href":61,"children":578},[60],[580],{"tag":553,"href":64,"children":581},[63],[583],{"tag":553,"href":67,"children":584},[66],[586],{"tag":553,"href":70,"children":587},[69],[589],{"tag":553,"href":73,"children":590},[72],[592],{"tag":553,"href":76,"children":593},[75],[595],{"tag":553,"href":79,"children":596},[78],[598],{"tag":553,"href":82,"children":599},[81],[601,666,726,791],{"slug":602,"published":5,"minutes":603,"category":7,"tags":604,"keywords":608,"about":618,"sources":625,"cover":660,"og":661,"expertise":85,"locales":662,"lang":87,"title":663,"description":664,"coverAlt":665},"openai-dots-always-on-agents-impact",14,[432,605,606,607],"AI agents","GPT-6 Astra","AI governance",[432,609,610,611,612,613,614,615,616,617],"what are OpenAI dots","OpenAI dots impact","always-on AI agents","GPT-6 Astra agents","dots Auto-review and Custom Rules","specialist dots for enterprise","OpenAI dots EU availability","OpenAI DevDay 2026","AI agents in the workplace",[619,621,624],{"name":453,"url":620},"https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FOpenAI",{"name":622,"url":623},"ChatGPT","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FChatGPT",{"name":26,"url":27},[626,629,632,633,636,639,642,645,648,651,654,657],{"title":627,"url":628},"OpenAI: Introducing dots (29 September 2026)","https:\u002F\u002Fopenai.com\u002Findex\u002Fintroducing-dots\u002F",{"title":630,"url":631},"OpenAI: How we build safety, security and privacy into dots","https:\u002F\u002Fopenai.com\u002Findex\u002Fhow-we-build-safety-security-and-privacy-into-dots\u002F",{"title":72,"url":73},{"title":634,"url":635},"TechCrunch: OpenAI launches Dots, its bubbly agentic avatar","https:\u002F\u002Ftechcrunch.com\u002F2026\u002F09\u002F29\u002Fopenai-launches-dots-its-bubbly-agentic-avatar\u002F",{"title":637,"url":638},"Unite.AI: OpenAI rolls out dots agents powered by GPT-6 Astra in ChatGPT","https:\u002F\u002Fwww.unite.ai\u002Fopenai-rolls-out-dots-agents-powered-by-gpt-6-astra-in-chatgpt\u002F",{"title":640,"url":641},"MediaNama: OpenAI launches dots that keep working without user prompts","https:\u002F\u002Fwww.medianama.com\u002F2026\u002F10\u002F223-openai-launches-dots-devday-2026\u002F",{"title":643,"url":644},"PYMNTS: OpenAI launches dots to capture AI agent market","https:\u002F\u002Fwww.pymnts.com\u002Fnews\u002Fartificial-intelligence\u002F2026\u002Fopenai-launches-dots-to-capture-ai-agent-market\u002F",{"title":646,"url":647},"Yahoo Finance: OpenAI debuts Dots AI agents in challenge to Meta's Muse","https:\u002F\u002Ffinance.yahoo.com\u002Ftechnology\u002Farticle\u002Fopenai-debuts-dots-ai-agents-in-challenge-to-metas-popular-muse-agent-174616593.html",{"title":649,"url":650},"CNBC: OpenAI abandons plan to release upcoming model as safety concerns escalate","https:\u002F\u002Fwww.cnbc.com\u002F2026\u002F09\u002F28\u002Fopenai-abandons-plan-to-release-upcoming-model-as-safety-concerns-escalate.html",{"title":652,"url":653},"The Hacker News: OpenAI shelves GPT-6.1 Astra after tests find deception and unauthorized actions","https:\u002F\u002Fthehackernews.com\u002F2026\u002F09\u002Fopenai-shelves-gpt-61-astra-after-tests.html",{"title":655,"url":656},"Al Jazeera: OpenAI launches dots, personal AI assistant built to handle everything","https:\u002F\u002Fwww.aljazeera.com\u002Feconomy\u002F2026\u002F9\u002F30\u002Fopenai-launches-dots-personal-ai-assistant-built-to-handle-everything",{"title":658,"url":659},"RedactSure: Do OpenAI dots Custom Rules control what the agent sees?","https:\u002F\u002Fredactsure.com\u002Fresearch\u002Fdo-openai-dots-custom-rules-control-what-the-agent-sees","\u002Fimages\u002Fblog\u002Fopenai-dots-always-on-agents-impact\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fopenai-dots-always-on-agents-impact\u002Fog.jpg",[87,88,89],"OpenAI dots: what always-on agents will change, and what they will not","OpenAI dots are always-on GPT-6 Astra agents with their own computer. What launched, how the safeguards work, and what changes for work, IT, SaaS and Europe.","Diagram: a dot running on GPT-6 Astra fans out to Slack and Teams, more than 4,000 apps, its own cloud computer, and a person who approves and reviews.",{"slug":667,"published":5,"minutes":6,"category":7,"tags":668,"keywords":672,"about":683,"sources":691,"cover":720,"og":721,"expertise":85,"locales":722,"lang":87,"title":723,"description":724,"coverAlt":725},"human-in-the-loop-ai-agents",[669,605,670,671],"Human in the loop","Approval gates","Agent safety",[673,674,675,676,677,678,679,680,681,682],"human in the loop AI agents","human in the loop KI-Agent","Mensch im Loop KI","AI agent approval gates","agent approval fatigue","LangGraph interrupt human in the loop","OpenAI Agents SDK needs_approval","Claude Agent SDK canUseTool","AI agent audit trail","risk tiers for AI agent actions",[684,687,688],{"name":685,"url":686},"Human-in-the-loop","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FHuman-in-the-loop",{"name":26,"url":27},{"name":689,"url":690},"Audit trail","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FAudit_trail",[692,695,698,701,704,707,710,713,716,719],{"title":693,"url":694},"LangChain docs: LangGraph interrupts","https:\u002F\u002Fdocs.langchain.com\u002Foss\u002Fpython\u002Flanggraph\u002Finterrupts",{"title":696,"url":697},"LangChain docs: Human-in-the-loop (HumanInTheLoopMiddleware)","https:\u002F\u002Fdocs.langchain.com\u002Foss\u002Fpython\u002Flangchain\u002Fhuman-in-the-loop",{"title":699,"url":700},"OpenAI Agents SDK (Python): Human in the loop","https:\u002F\u002Fgithub.com\u002Fopenai\u002Fopenai-agents-python\u002Fblob\u002Fmain\u002Fdocs\u002Fhuman_in_the_loop.md",{"title":702,"url":703},"Claude Agent SDK: Handle approvals and user input","https:\u002F\u002Fcode.claude.com\u002Fdocs\u002Fen\u002Fagent-sdk\u002Fuser-input",{"title":705,"url":706},"Claude Agent SDK: Configure permissions","https:\u002F\u002Fcode.claude.com\u002Fdocs\u002Fen\u002Fagent-sdk\u002Fpermissions",{"title":708,"url":709},"Claude Code docs: Hooks (PreToolUse defer, PermissionRequest)","https:\u002F\u002Fcode.claude.com\u002Fdocs\u002Fen\u002Fhooks",{"title":711,"url":712},"Anthropic Engineering: Claude Code auto mode","https:\u002F\u002Fanthropic.com\u002Fengineering\u002Fclaude-code-auto-mode",{"title":714,"url":715},"DevOps.com: Anthropic makes Claude Code auto mode the default","https:\u002F\u002Fdevops.com\u002Fanthropic-makes-claude-codes-auto-mode-the-default-betting-automation-beats-manual-review\u002F",{"title":717,"url":718},"EU AI Act, Article 14: Human oversight","https:\u002F\u002Fartificialintelligenceact.eu\u002Farticle\u002F14\u002F",{"title":630,"url":631},"\u002Fimages\u002Fblog\u002Fhuman-in-the-loop-ai-agents\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fhuman-in-the-loop-ai-agents\u002Fog.jpg",[87,88,89],"Human in the loop for AI agents: where to put approval gates","Where approval gates belong in an AI agent, how to avoid rubber-stamping, and how interrupt and resume work in LangGraph and the OpenAI and Claude agent SDKs.","Diagram: an agent proposes an action, a risk gate sends it to automatic execution, to a human approval, or to a block, and every decision lands in an audit log.",{"slug":727,"published":5,"minutes":728,"category":7,"tags":729,"keywords":732,"about":743,"sources":751,"cover":785,"og":786,"expertise":85,"locales":787,"lang":87,"title":788,"description":789,"coverAlt":790},"multi-agent-systems-when-worth-it",12,[730,605,731,10],"Multi-agent systems","Orchestrator-worker",[733,734,735,736,737,738,739,740,741,742],"multi-agent systems when worth it","multi-agent vs single agent","orchestrator-worker pattern","multi-agent LLM token cost","why multi-agent systems fail","don't build multi-agents","subagents context isolation","multi-agent debate worth it","agent handoffs vs subagents","when to use multiple AI agents",[744,747,748],{"name":745,"url":746},"Multi-agent system","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FMulti-agent_system",{"name":26,"url":27},{"name":749,"url":750},"Large language model","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FLarge_language_model",[752,755,758,761,764,767,770,773,776,779,782],{"title":753,"url":754},"Anthropic Engineering: How we built our multi-agent research system","https:\u002F\u002Fwww.anthropic.com\u002Fengineering\u002Fmulti-agent-research-system",{"title":756,"url":757},"Anthropic: Building effective agents","https:\u002F\u002Fwww.anthropic.com\u002Fengineering\u002Fbuilding-effective-agents",{"title":759,"url":760},"Cognition: Don't Build Multi-Agents (12 June 2025)","https:\u002F\u002Fcognition.com\u002Fblog\u002Fdont-build-multi-agents",{"title":762,"url":763},"Cognition: Multi-Agents: What's Actually Working (22 April 2026)","https:\u002F\u002Fcognition.com\u002Fblog\u002Fmulti-agents-working",{"title":765,"url":766},"Google Research: Towards a science of scaling agent systems","https:\u002F\u002Fresearch.google\u002Fblog\u002Ftowards-a-science-of-scaling-agent-systems-when-and-why-agent-systems-work\u002F",{"title":768,"url":769},"arXiv 2512.08296: Towards a Science of Scaling Agent Systems","https:\u002F\u002Farxiv.org\u002Fabs\u002F2512.08296",{"title":771,"url":772},"arXiv 2503.13657: Why Do Multi-Agent LLM Systems Fail? (MAST)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2503.13657",{"title":774,"url":775},"arXiv 2305.14325: Improving Factuality and Reasoning in Language Models through Multiagent Debate","https:\u002F\u002Farxiv.org\u002Fabs\u002F2305.14325",{"title":777,"url":778},"arXiv 2502.08788: Stop Overvaluing Multi-Agent Debate","https:\u002F\u002Farxiv.org\u002Fabs\u002F2502.08788",{"title":780,"url":781},"OpenAI Agents SDK: Handoffs","https:\u002F\u002Fopenai.github.io\u002Fopenai-agents-python\u002Fhandoffs\u002F",{"title":783,"url":784},"Claude Code documentation: Subagents","https:\u002F\u002Fcode.claude.com\u002Fdocs\u002Fen\u002Fsub-agents","\u002Fimages\u002Fblog\u002Fmulti-agent-systems-when-worth-it\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fmulti-agent-systems-when-worth-it\u002Fog.jpg",[87,88,89],"Multi-agent systems: when they beat one agent, and when they do not","Orchestrator-worker, fan-out, critic, handoff: what multi-agent systems really buy you, what they cost in tokens, how they fail, and a table to decide.","Diagram: a lead agent fans out to four worker agents, each with its own isolated context window, and gathers their summaries back.",{"slug":792,"published":5,"minutes":6,"category":7,"tags":793,"keywords":799,"about":810,"sources":820,"cover":875,"og":876,"expertise":85,"locales":877,"lang":87,"title":878,"description":879,"coverAlt":880},"voice-agents-realtime-latency",[794,795,796,797,798],"Voice agents","Realtime API","Latency","Telephony","AI Act",[800,801,802,803,804,805,806,807,808,809],"voice agents","speech-to-speech vs STT LLM TTS","OpenAI Realtime API","Gemini Live API","voice agent latency budget","turn detection and barge-in","Pipecat vs LiveKit","AI voice agent SIP telephony","German Hungarian voice AI","AI Act Article 50 voice bot disclosure",[811,814,817],{"name":812,"url":813},"Voice user interface","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FVoice_user_interface",{"name":815,"url":816},"Speech recognition","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FSpeech_recognition",{"name":818,"url":819},"Artificial Intelligence Act","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FArtificial_Intelligence_Act",[821,824,827,830,833,836,839,842,845,848,851,854,857,860,863,866,869,872],{"title":822,"url":823},"OpenAI: Voice agents guide","https:\u002F\u002Fdevelopers.openai.com\u002Fapi\u002Fdocs\u002Fguides\u002Fvoice-agents",{"title":825,"url":826},"OpenAI: gpt-realtime model page","https:\u002F\u002Fdevelopers.openai.com\u002Fapi\u002Fdocs\u002Fmodels\u002Fgpt-realtime",{"title":828,"url":829},"OpenAI: Voice activity detection in the Realtime API","https:\u002F\u002Fdevelopers.openai.com\u002Fapi\u002Fdocs\u002Fguides\u002Frealtime-vad",{"title":831,"url":832},"OpenAI: Realtime API with SIP","https:\u002F\u002Fdevelopers.openai.com\u002Fapi\u002Fdocs\u002Fguides\u002Frealtime-sip",{"title":834,"url":835},"OpenAI: Realtime conversations (function calling, interruption)","https:\u002F\u002Fdevelopers.openai.com\u002Fapi\u002Fdocs\u002Fguides\u002Frealtime-conversations",{"title":837,"url":838},"Google: Gemini Live API overview","https:\u002F\u002Fai.google.dev\u002Fgemini-api\u002Fdocs\u002Flive",{"title":840,"url":841},"Google: Gemini Live API capabilities guide","https:\u002F\u002Fai.google.dev\u002Fgemini-api\u002Fdocs\u002Flive-guide",{"title":843,"url":844},"Deepgram: Flux quickstart","https:\u002F\u002Fdevelopers.deepgram.com\u002Fdocs\u002Fflux\u002Fquickstart",{"title":846,"url":847},"Deepgram: Models and languages overview","https:\u002F\u002Fdevelopers.deepgram.com\u002Fdocs\u002Fmodels-languages-overview",{"title":849,"url":850},"ElevenLabs: Agents platform overview","https:\u002F\u002Felevenlabs.io\u002Fdocs\u002Feleven-agents\u002Foverview",{"title":852,"url":853},"ElevenLabs: Text to speech models and languages","https:\u002F\u002Felevenlabs.io\u002Fdocs\u002Foverview\u002Fcapabilities\u002Ftext-to-speech",{"title":855,"url":856},"LiveKit: Agents overview","https:\u002F\u002Fdocs.livekit.io\u002Fagents\u002F",{"title":858,"url":859},"LiveKit: Turn detector","https:\u002F\u002Fdocs.livekit.io\u002Fagents\u002Flogic\u002Fturns\u002Fturn-detector\u002F",{"title":861,"url":862},"Pipecat: Introduction","https:\u002F\u002Fdocs.pipecat.ai\u002Fgetting-started\u002Fintroduction",{"title":864,"url":865},"Pipecat: Smart Turn model (GitHub)","https:\u002F\u002Fgithub.com\u002Fpipecat-ai\u002Fsmart-turn",{"title":867,"url":868},"Fora Soft: Voice AI agents on LiveKit, 2026 engineer playbook","https:\u002F\u002Fwww.forasoft.com\u002Fblog\u002Farticle\u002Fvoice-ai-agents-livekit-guide",{"title":870,"url":871},"EU AI Act: Article 50, transparency obligations","https:\u002F\u002Fartificialintelligenceact.eu\u002Farticle\u002F50\u002F",{"title":873,"url":874},"Jones Walker: Yes, August 2 still matters (AI Act delay and Article 50)","https:\u002F\u002Fwww.joneswalker.com\u002Fen\u002Finsights\u002Fblogs\u002Fai-law-blog\u002Fyes-august-2-still-matters-the-eu-approved-a-high-risk-ai-delay-but-most-trans.html?id=102nbon","\u002Fimages\u002Fblog\u002Fvoice-agents-realtime-latency\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fvoice-agents-realtime-latency\u002Fog.jpg",[87,88,89],"Building voice agents: realtime speech-to-speech or STT, LLM and TTS?","Realtime speech-to-speech or a cascaded pipeline? Latency budget per stage, turn-taking, tool calls, SIP, German and Hungarian quality, and AI Act disclosure.","Diagram: a caller reaches a voice agent over SIP or WebRTC, which fans out to turn detection, speech recognition, an LLM with tools and speech synthesis.",1791009037084]