[{"data":1,"prerenderedAt":1317},["ShallowReactive",2],{"blog-token-saving-tools-coding-agents-top-20-en":3},{"slug":4,"published":5,"minutes":6,"category":7,"tags":8,"keywords":13,"about":25,"sources":35,"cover":96,"og":97,"expertise":98,"locales":99,"lang":100,"title":103,"description":104,"coverAlt":105,"metaTitle":106,"takeaways":107,"faq":113,"toc":126,"blocks":163,"others":1088},"token-saving-tools-coding-agents-top-20","2026-09-28",17,"agents",[9,10,11,12],"Claude Code","token usage","context engineering","developer tools",[14,15,16,17,18,19,20,21,22,23,24,11],"reduce Claude Code token usage","token saving tools for coding agents","rtk token killer","lean-ctx","context-mode MCP","Serena MCP","Repomix compress","MCP tool search","prompt caching","ccusage","coding agent cost",[26,29,32],{"name":27,"url":28},"Large language model","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FLarge_language_model",{"name":30,"url":31},"Claude (language model)","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FClaude_(language_model)",{"name":33,"url":34},"Prompt engineering","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FPrompt_engineering",[36,39,42,45,48,51,54,57,60,63,66,69,72,75,78,81,84,87,90,93],{"title":37,"url":38},"rtk: a CLI proxy that filters shell output for coding agents (GitHub)","https:\u002F\u002Fgithub.com\u002Frtk-ai\u002Frtk",{"title":40,"url":41},"lean-ctx: context read modes and shell compression for coding agents (GitHub)","https:\u002F\u002Fgithub.com\u002Fyvgude\u002Flean-ctx",{"title":43,"url":44},"context-mode: sandboxed tool output with SQLite FTS5 search (GitHub)","https:\u002F\u002Fgithub.com\u002Fmksglu\u002Fcontext-mode",{"title":46,"url":47},"Serena: semantic code retrieval and editing over MCP (GitHub)","https:\u002F\u002Fgithub.com\u002Foraios\u002Fserena",{"title":49,"url":50},"token-savior: symbol index, memory and bash compaction over MCP (GitHub)","https:\u002F\u002Fgithub.com\u002FMibayy\u002Ftoken-savior",{"title":52,"url":53},"code-review-graph: a code graph for blast-radius reviews (GitHub)","https:\u002F\u002Fgithub.com\u002Ftirth8205\u002Fcode-review-graph",{"title":55,"url":56},"Aider documentation: repository map","https:\u002F\u002Faider.chat\u002Fdocs\u002Frepomap.html",{"title":58,"url":59},"Repomix: pack a repository into one AI-friendly file (GitHub)","https:\u002F\u002Fgithub.com\u002Fyamadashy\u002Frepomix",{"title":61,"url":62},"Context7: up-to-date library documentation for LLMs (GitHub)","https:\u002F\u002Fgithub.com\u002Fupstash\u002Fcontext7",{"title":64,"url":65},"claude-context: hybrid code search MCP (GitHub)","https:\u002F\u002Fgithub.com\u002Fzilliztech\u002Fclaude-context",{"title":67,"url":68},"caveman: terse output modes for coding agents (GitHub)","https:\u002F\u002Fgithub.com\u002FJuliusBrussee\u002Fcaveman",{"title":70,"url":71},"claude-token-efficient: an eight-rule CLAUDE.md (GitHub)","https:\u002F\u002Fgithub.com\u002Fdrona23\u002Fclaude-token-efficient",{"title":73,"url":74},"claude-code-router: a local model gateway for coding agents (GitHub)","https:\u002F\u002Fgithub.com\u002Fmusistudio\u002Fclaude-code-router",{"title":76,"url":77},"ccusage: token and cost reports from local agent logs (GitHub)","https:\u002F\u002Fgithub.com\u002Fryoppippi\u002Fccusage",{"title":79,"url":80},"LLMLingua: prompt compression (Microsoft, GitHub)","https:\u002F\u002Fgithub.com\u002Fmicrosoft\u002FLLMLingua",{"title":82,"url":83},"ComputingForGeeks: tools that reduce Claude Code token usage, tested (April 2026)","https:\u002F\u002Fcomputingforgeeks.com\u002Freduce-claude-code-token-usage-tools\u002F",{"title":85,"url":86},"Claude Code docs: manage costs effectively","https:\u002F\u002Fcode.claude.com\u002Fdocs\u002Fen\u002Fcosts",{"title":88,"url":89},"Claude Code docs: MCP, output limits and tool search","https:\u002F\u002Fcode.claude.com\u002Fdocs\u002Fen\u002Fmcp",{"title":91,"url":92},"Claude API docs: tool search tool","https:\u002F\u002Fplatform.claude.com\u002Fdocs\u002Fen\u002Fagents-and-tools\u002Ftool-use\u002Ftool-search-tool",{"title":94,"url":95},"Claude API docs: prompt caching","https:\u002F\u002Fplatform.claude.com\u002Fdocs\u002Fen\u002Fbuild-with-claude\u002Fprompt-caching","\u002Fimages\u002Fblog\u002Ftoken-saving-tools-coding-agents-top-20\u002Fcover.webp","\u002Fimages\u002Fblog\u002Ftoken-saving-tools-coding-agents-top-20\u002Fog.jpg","ai-engineer",[100,101,102],"en","de","hu","Top 20 ways to cut coding-agent tokens: rtk, lean-ctx, Serena and more, ranked by evidence","rtk, lean-ctx, context-mode, Serena and 16 more token savers for coding agents, ranked by evidence, with my own measurements on a real Nuxt codebase.","Bar chart falling from a 15,100-token build log to about 370 tokens after filtering, under the heading Top 20 token savers.","Top 20 token-saving tools for coding agents · Balázs Csorba",[108,109,110,111,112],"Most token-saver percentages are output reduction on the commands a tool is best at, not bill reduction, so treat them as upper bounds.","The free built-in habits come first: measure with \u002Fusage and ccusage, \u002Fclear between tasks, keep the cached prefix stable, trim MCP servers and lower effort.","Shell output was the biggest leak I measured: a 15,100-token build log shrank to about 370 tokens, with every useful line kept, by dropping colour codes, warnings and route lists.","On the read side, symbol navigation (code intelligence plugins, Serena, lean-ctx) beats reading whole files. Repomix --compress cut TypeScript by 79% in my run but made Vue files 30% larger.","The one independent head-to-head test I found put token-savior, claude-token-efficient and caveman in front at 38–43%, on a single repository.",[114,117,120,123],{"q":115,"a":116},"What is the best tool to reduce Claude Code token usage?","There is no single one, because each tool works on a different part of a turn. Start with the free built-ins: \u002Fclear between tasks, a stable prefix for caching, fewer MCP servers and lower effort. Then add one output filter such as rtk or context-mode and one read-side tool such as a code intelligence plugin or Serena, and measure before and after.",{"q":118,"a":119},"Does rtk really save 90% of tokens?","On noisy commands such as test runs and build logs it removes 60–90% of the output, and an independent test confirmed that range. On commands whose output is already short it saves close to nothing, and it does not touch Claude Code’s built-in Read, Grep and Glob tools, so the effect on a whole session is much smaller than the headline.",{"q":121,"a":122},"Is it safe to compress context for a coding agent?","Methods that keep structure are safe: filtering noise out of logs, reading signatures before bodies and looking up symbols through a language server. Lossy prompt compression that drops individual tokens, such as LLMLingua, suits prose and retrieval, but not code the agent has to edit exactly.",{"q":124,"a":125},"How do I measure my token usage in Claude Code?","Use \u002Fusage for the current session, including its prompt cache hit rate, and \u002Fcontext to see what fills the context window. For history across sessions, ccusage reads the local logs and reports usage by day, month, session or five-hour block.",[127,130,133,136,139,142,145,148,151,154,157,160],{"id":128,"title":129},"where-tokens-go","Where the tokens in an agent turn come from",{"id":131,"title":132},"measured","What I measured on this site",{"id":134,"title":135},"ranking","The top 20, ranked",{"id":137,"title":138},"built-in","Ranks 1–5: built-in habits that cost nothing",{"id":140,"title":141},"output-filters","Ranks 6–8: filter tool output before the model reads it",{"id":143,"title":144},"code-navigation","Ranks 9–11: read symbols, not whole files",{"id":146,"title":147},"context-hygiene","Ranks 12–13: keep the main context small",{"id":149,"title":150},"indexes","Ranks 14–18: indexes, maps and documentation",{"id":152,"title":153},"output-and-routing","Ranks 19–20: output style and routing",{"id":155,"title":156},"skip","What I would skip, or use with care",{"id":158,"title":159},"starter-stack","A starter stack for one afternoon",{"id":161,"title":162},"sources","Sources",[164,168,182,204,207,226,235,247,248,259,266,277,288,344,347,357,358,361,579,586,648,649,652,667,669,686,688,691,693,727,729,749,750,752,775,777,782,784,787,789,815,816,818,827,829,851,853,859,860,862,870,872,875,876,878,883,885,899,901,915,917,930,932,937,938,940,949,951,959,960,988,989,1017,1025,1026],{"type":165,"content":166},"paragraph",[167],"Most advice on cutting coding-agent costs comes down to a screenshot of a counter that says 90% saved. I wanted to know which of these tools hold up, so I went through the READMEs and benchmarks of more than twenty token savers, compared them with the one independent head-to-head test I could find, and measured two of the ideas on this site’s own codebase. This is the ranking that came out of it.",{"type":165,"content":169},[170,171,176,177,181],"The ranking is written for Claude Code, because that is where I work and where the official documentation is most specific, but most tools on the list also plug into Cursor, Codex, OpenCode or any other agent that speaks MCP or runs shell hooks. For the background on why context size drives both cost and quality, start with ",{"tag":172,"to":173,"children":174},"link","\u002Fblog\u002Fagents-md-skills-mcp-cli-decision-matrix",[175],"context engineering for coding agents"," and ",{"tag":172,"to":178,"children":179},"\u002Fblog\u002Fllm-cost-latency-prompt-caching-routing",[180],"prompt caching and model routing",".",{"type":183,"variant":184,"title":185,"body":186},"callout","warn","Read every percentage twice",[187,194],[188,189,193],"A tool that removes 90% of ",{"tag":190,"children":191},"code",[192],"git diff"," output has not cut your bill by 90%. Tool output is one slice of a turn; the cached prefix, the conversation history and the model’s own output are the rest. Most READMEs report output reduction on the commands the tool is best at, often with a bytes-divided-by-four token estimate.",[195,196,200,201,203],"In this post, ",{"tag":197,"children":198},"strong",[199],"claimed"," means the project’s own number and ",{"tag":197,"children":202},[131]," means an independent test or my own run. Where I know which tokenizer or estimate was used, I say so.",{"type":205,"level":206,"id":128,"text":129},"heading",2,{"type":165,"content":208},[209,210,213,214,217,218,221,222,225],"Every request an agent sends carries four kinds of tokens, and each tool on this list works on one of them. The ",{"tag":197,"children":211},[212],"prefix"," is the system prompt, the tool definitions and CLAUDE.md. ",{"tag":197,"children":215},[216],"Reads"," are the files and search results the agent pulls in to understand the code. ",{"tag":197,"children":219},[220],"Tool output"," is what shell commands, test runners and MCP servers print back. ",{"tag":197,"children":223},[224],"Model output"," is the thinking and the answer. All four pile up in the history, which is sent again on every turn.",{"type":227,"attrs":228,"inner":232,"caption":233},"diagram",{"viewBox":229,"role":230,"aria-labelledby":231},"0 0 720 372","img","d1-tok-t d1-tok-d","\u003Ctitle id=\"d1-tok-t\">Where a turn’s tokens come from\u003C\u002Ftitle>\u003Cdesc id=\"d1-tok-d\">Four columns. Prefix: tool definitions and CLAUDE.md, cut by tool search, fewer MCP servers, CLIs instead of MCP, a lean CLAUDE.md and a stable prefix for caching. Reads: files and search results, cut by code intelligence, Serena and lean-ctx, token-savior, repo maps and Repomix, and claude-context. Tool output: shell, tests and MCP, cut by rtk, context-mode, your own hooks, a cap on MCP output and subagents. Model output: thinking and answers, cut by effort, model choice, Haiku for subagents, caveman and terse style rules. Below them a bar: the history, re-sent every turn, kept small with clear, compact and subagents and kept cheap by a stable cached prefix.\u003C\u002Fdesc>\u003Ctext x=\"20\" y=\"28\" class=\"d-title\">Where a turn’s tokens come from\u003C\u002Ftext>\u003Ctext x=\"700\" y=\"28\" text-anchor=\"end\" class=\"d-label\">and what cuts them\u003C\u002Ftext>\u003Crect x=\"20\" y=\"46\" width=\"164\" height=\"64\" rx=\"10\" class=\"d-box\" \u002F>\u003Ctext x=\"102\" y=\"73\" text-anchor=\"middle\" class=\"d-text\">Prefix\u003C\u002Ftext>\u003Ctext x=\"102\" y=\"95\" text-anchor=\"middle\" class=\"d-small\">tools, CLAUDE.md\u003C\u002Ftext>\u003Cpath d=\"M102 110 V124\" class=\"d-line\" \u002F>\u003Crect x=\"20\" y=\"124\" width=\"164\" height=\"150\" rx=\"10\" class=\"d-box d-dash\" \u002F>\u003Ctext x=\"102\" y=\"152\" text-anchor=\"middle\" class=\"d-small\">tool search\u003C\u002Ftext>\u003Ctext x=\"102\" y=\"178\" text-anchor=\"middle\" class=\"d-small\">fewer MCP servers\u003C\u002Ftext>\u003Ctext x=\"102\" y=\"204\" text-anchor=\"middle\" class=\"d-small\">CLI over MCP\u003C\u002Ftext>\u003Ctext x=\"102\" y=\"230\" text-anchor=\"middle\" class=\"d-small\">lean CLAUDE.md\u003C\u002Ftext>\u003Ctext x=\"102\" y=\"256\" text-anchor=\"middle\" class=\"d-small\">stable for caching\u003C\u002Ftext>\u003Crect x=\"192\" y=\"46\" width=\"164\" height=\"64\" rx=\"10\" class=\"d-sky\" \u002F>\u003Ctext x=\"274\" y=\"73\" text-anchor=\"middle\" class=\"d-text\">Reads\u003C\u002Ftext>\u003Ctext x=\"274\" y=\"95\" text-anchor=\"middle\" class=\"d-small\">files, search results\u003C\u002Ftext>\u003Cpath d=\"M274 110 V124\" class=\"d-line\" \u002F>\u003Crect x=\"192\" y=\"124\" width=\"164\" height=\"150\" rx=\"10\" class=\"d-box d-dash\" \u002F>\u003Ctext x=\"274\" y=\"152\" text-anchor=\"middle\" class=\"d-small\">code intelligence\u003C\u002Ftext>\u003Ctext x=\"274\" y=\"178\" text-anchor=\"middle\" class=\"d-small\">Serena, lean-ctx\u003C\u002Ftext>\u003Ctext x=\"274\" y=\"204\" text-anchor=\"middle\" class=\"d-small\">token-savior\u003C\u002Ftext>\u003Ctext x=\"274\" y=\"230\" text-anchor=\"middle\" class=\"d-small\">repo maps, Repomix\u003C\u002Ftext>\u003Ctext x=\"274\" y=\"256\" text-anchor=\"middle\" class=\"d-small\">claude-context\u003C\u002Ftext>\u003Crect x=\"364\" y=\"46\" width=\"164\" height=\"64\" rx=\"10\" class=\"d-accent\" \u002F>\u003Ctext x=\"446\" y=\"73\" text-anchor=\"middle\" class=\"d-text\">Tool output\u003C\u002Ftext>\u003Ctext x=\"446\" y=\"95\" text-anchor=\"middle\" class=\"d-small\">shell, tests, MCP\u003C\u002Ftext>\u003Cpath d=\"M446 110 V124\" class=\"d-line\" \u002F>\u003Crect x=\"364\" y=\"124\" width=\"164\" height=\"150\" rx=\"10\" class=\"d-box d-dash\" \u002F>\u003Ctext x=\"446\" y=\"152\" text-anchor=\"middle\" class=\"d-small\">rtk\u003C\u002Ftext>\u003Ctext x=\"446\" y=\"178\" text-anchor=\"middle\" class=\"d-small\">context-mode\u003C\u002Ftext>\u003Ctext x=\"446\" y=\"204\" text-anchor=\"middle\" class=\"d-small\">your own hooks\u003C\u002Ftext>\u003Ctext x=\"446\" y=\"230\" text-anchor=\"middle\" class=\"d-small\">MCP output cap\u003C\u002Ftext>\u003Ctext x=\"446\" y=\"256\" text-anchor=\"middle\" class=\"d-small\">subagents\u003C\u002Ftext>\u003Crect x=\"536\" y=\"46\" width=\"164\" height=\"64\" rx=\"10\" class=\"d-gold\" \u002F>\u003Ctext x=\"618\" y=\"73\" text-anchor=\"middle\" class=\"d-text\">Model output\u003C\u002Ftext>\u003Ctext x=\"618\" y=\"95\" text-anchor=\"middle\" class=\"d-small\">thinking, answers\u003C\u002Ftext>\u003Cpath d=\"M618 110 V124\" class=\"d-line\" \u002F>\u003Crect x=\"536\" y=\"124\" width=\"164\" height=\"150\" rx=\"10\" class=\"d-box d-dash\" \u002F>\u003Ctext x=\"618\" y=\"152\" text-anchor=\"middle\" class=\"d-small\">\u002Feffort\u003C\u002Ftext>\u003Ctext x=\"618\" y=\"178\" text-anchor=\"middle\" class=\"d-small\">\u002Fmodel\u003C\u002Ftext>\u003Ctext x=\"618\" y=\"204\" text-anchor=\"middle\" class=\"d-small\">Haiku for subagents\u003C\u002Ftext>\u003Ctext x=\"618\" y=\"230\" text-anchor=\"middle\" class=\"d-small\">caveman\u003C\u002Ftext>\u003Ctext x=\"618\" y=\"256\" text-anchor=\"middle\" class=\"d-small\">terse style rules\u003C\u002Ftext>\u003Crect x=\"20\" y=\"292\" width=\"680\" height=\"64\" rx=\"10\" class=\"d-mint\" \u002F>\u003Ctext x=\"360\" y=\"319\" text-anchor=\"middle\" class=\"d-text\">History: all of the above is sent again on every turn\u003C\u002Ftext>\u003Ctext x=\"360\" y=\"341\" text-anchor=\"middle\" class=\"d-small\">\u002Fclear · \u002Fcompact · subagents · a stable prefix keeps it cached\u003C\u002Ftext>",[234],"Each tool works on one of four token sources; the history multiplies all of them by the number of turns.",{"type":165,"content":236},[237,238,242,243,246],"The prices are not symmetric, and that decides where savings matter. On Claude Opus 5.5 a cached input token costs $0.20 per million, a fresh one $4 and an output token $20, according to the ",{"tag":239,"href":95,"children":240},"a",[241],"prompt caching docs",". A stable, cached prefix is almost free after the first request. What costs money is new content: fresh reads, fresh tool output and everything the model writes, thinking included. Claude Code’s own ",{"tag":239,"href":86,"children":244},[245],"cost guide"," puts the average at about $13 per developer per active day, and below $30 for 90% of users, so the problem is rarely one big bill. It is a steady leak across many turns.",{"type":205,"level":206,"id":131,"text":132},{"type":165,"content":249},[250,251,254,255,258],"Two of the claims were cheap to test here. This site is a Nuxt 4 app with about 110 Vue, TypeScript and script files, and its build prints a lot. I ran ",{"tag":239,"href":59,"children":252},[253],"Repomix"," 1.18.1 with its default o200k_base tokenizer on the source, once plain and once with ",{"tag":190,"children":256},[257],"--compress",", which uses tree-sitter to keep signatures and drop function bodies. The README puts the reduction at about 70%.",{"type":227,"attrs":260,"inner":263,"caption":264},{"viewBox":261,"role":230,"aria-labelledby":262},"0 0 720 324","d2-tok-t d2-tok-d","\u003Ctitle id=\"d2-tok-t\">Repomix 1.18.1 on this site\u003C\u002Ftitle>\u003Cdesc id=\"d2-tok-d\">Grouped bars of token counts. TypeScript and script files, 62 files: 133,146 tokens plain, 28,402 with compress, 79% fewer. Vue components, 49 files: 72,298 plain, 94,096 with compress, 30% more. Whole source, 111 files: 205,115 plain, 122,132 with compress, 40% fewer.\u003C\u002Fdesc>\u003Ctext x=\"20\" y=\"28\" class=\"d-title\">Repomix 1.18.1 on this site\u003C\u002Ftext>\u003Ctext x=\"700\" y=\"28\" text-anchor=\"end\" class=\"d-label\">tokens, o200k_base\u003C\u002Ftext>\u003Ctext x=\"20\" y=\"76\" class=\"d-text\">TypeScript, scripts\u003C\u002Ftext>\u003Ctext x=\"20\" y=\"96\" class=\"d-small\">62 files\u003C\u002Ftext>\u003Crect x=\"230\" y=\"50\" width=\"259.7\" height=\"26\" rx=\"6\" class=\"d-box\" \u002F>\u003Ctext x=\"499.7\" y=\"68\" class=\"d-label\">133,146\u003C\u002Ftext>\u003Crect x=\"230\" y=\"82\" width=\"55.4\" height=\"26\" rx=\"6\" class=\"d-accent\" \u002F>\u003Ctext x=\"295.4\" y=\"100\" class=\"d-label\">28,402  −79%\u003C\u002Ftext>\u003Ctext x=\"20\" y=\"156\" class=\"d-text\">Vue components\u003C\u002Ftext>\u003Ctext x=\"20\" y=\"176\" class=\"d-small\">49 files\u003C\u002Ftext>\u003Crect x=\"230\" y=\"130\" width=\"141.0\" height=\"26\" rx=\"6\" class=\"d-box\" \u002F>\u003Ctext x=\"381.0\" y=\"148\" class=\"d-label\">72,298\u003C\u002Ftext>\u003Crect x=\"230\" y=\"162\" width=\"183.5\" height=\"26\" rx=\"6\" class=\"d-gold\" \u002F>\u003Ctext x=\"423.5\" y=\"180\" class=\"d-label\">94,096  +30%\u003C\u002Ftext>\u003Ctext x=\"20\" y=\"236\" class=\"d-text\">Whole source\u003C\u002Ftext>\u003Ctext x=\"20\" y=\"256\" class=\"d-small\">111 files\u003C\u002Ftext>\u003Crect x=\"230\" y=\"210\" width=\"400.0\" height=\"26\" rx=\"6\" class=\"d-box\" \u002F>\u003Ctext x=\"640.0\" y=\"228\" class=\"d-label\">205,115\u003C\u002Ftext>\u003Crect x=\"230\" y=\"242\" width=\"238.2\" height=\"26\" rx=\"6\" class=\"d-accent\" \u002F>\u003Ctext x=\"478.2\" y=\"260\" class=\"d-label\">122,132  −40%\u003C\u002Ftext>\u003Crect x=\"230\" y=\"294\" width=\"16\" height=\"12\" rx=\"3\" class=\"d-box\" \u002F>\u003Ctext x=\"254\" y=\"305\" class=\"d-small\">plain pack\u003C\u002Ftext>\u003Crect x=\"380\" y=\"294\" width=\"16\" height=\"12\" rx=\"3\" class=\"d-accent\" \u002F>\u003Ctext x=\"404\" y=\"305\" class=\"d-small\">--compress\u003C\u002Ftext>\u003Crect x=\"530\" y=\"294\" width=\"16\" height=\"12\" rx=\"3\" class=\"d-gold\" \u002F>\u003Ctext x=\"554\" y=\"305\" class=\"d-small\">--compress, larger\u003C\u002Ftext>",[265],"Measured on this site’s source: compression works on TypeScript and backfires on Vue single-file components.",{"type":165,"content":267},[268,269,272,273,276],"On the TypeScript and script files the claim holds and then some: 133,146 tokens became 28,402, a 79% cut. On the Vue single-file components it went the other way: 72,298 tokens became 94,096, 30% ",{"tag":197,"children":270},[271],"more",", and 46 of the 49 files grew. Looking at the output, the compressed Vue files kept the full template and then repeated fragments of it as extra chunks between ",{"tag":190,"children":274},[275],"⋮----"," markers. Across the whole codebase the saving was 40%, not 70%.",{"type":165,"content":278},[279,280,283,284,287],"The second test was the use case behind ",{"tag":239,"href":38,"children":281},[282],"rtk"," and every other output filter on this list: a noisy command. A full ",{"tag":190,"children":285},[286],"npm run generate"," of this site writes 501 lines. I stripped the log in stages and estimated tokens as bytes divided by four, the same rough heuristic rtk uses.",{"type":289,"head":290,"rows":299},"table",[291,293,295,297],[292],"Stage",[294],"Bytes",[296],"Tokens (bytes\u002F4)",[298],"Lines",[300,309,317,326,335],[301,303,305,307],[302],"Raw log, as written to a file",[304],"60,311",[306],"~15,100",[308],"501",[310,312,314,316],[311],"Without ANSI colour codes",[313],"44,426",[315],"~11,100",[308],[318,320,322,324],[319],"Without Node experimental warnings",[321],"42,144",[323],"~10,500",[325],"481",[327,329,331,333],[328],"Without the per-route and per-chunk lists",[330],"1,479",[332],"~370",[334],"46",[336,338,340,342],[337],"Only summary and error lines",[339],"620",[341],"~155",[343],"12",{"type":165,"content":345},[346],"Colour codes alone were a quarter of the log, even though it was redirected to a file. The lists of prerendered routes and built chunks were 96% of what remained, and on a green build none of it tells a model anything. Dropping those three things keeps every line a person would act on and removes 97.5% of the tokens. On a red build you need the error lines and a few lines around them, which is exactly what a good filter keeps.",{"type":183,"variant":348,"title":349,"body":350},"tip","The lesson",[351],[352,353,356],"Both results point the same way. The savings are real, but they depend on your stack and your commands, not on the number in the README. Measure one real session before and after, with ",{"tag":190,"children":354},[355],"\u002Fusage"," or ccusage, before you keep a tool installed.",{"type":205,"level":206,"id":134,"text":135},{"type":165,"content":359},[360],"I ranked by four things, in this order: whether the saving is backed by something other than the author’s own benchmark, how much of a real session it touches, how long setup takes, and what it risks, from lossy output to a restrictive license. Built-in habits come first because they are free and documented; third-party tools follow, ordered by evidence.",{"type":289,"head":362,"rows":373},[363,365,367,369,371],[364],"#",[366],"Tool or habit",[368],"Layer",[370],"Evidence",[372],"Setup",[374,391,408,417,426,435,445,455,468,477,487,495,503,511,520,530,541,550,560,569],[375,377,385,387,389],[376],"1",[378,380,381,384],{"tag":190,"children":379},[355],", ",{"tag":190,"children":382},[383],"\u002Fcontext",", ccusage",[386],"all",[388],"the measurement itself",[390],"minutes",[392,394,402,404,406],[393],"2",[395,398,399],{"tag":190,"children":396},[397],"\u002Fclear"," and a focused ",{"tag":190,"children":400},[401],"\u002Fcompact",[403],"history",[405],"official docs",[407],"none",[409,411,413,414,416],[410],"3",[412],"A stable prefix for prompt caching",[212],[415],"official pricing",[407],[418,420,422,423,425],[419],"4",[421],"Tool search, fewer MCP servers, CLIs",[212],[424],"official: over 85% of tool definitions",[390],[427,429,431,433,434],[428],"5",[430],"Effort and model choice",[432],"model output",[405],[407],[436,438,439,441,443],[437],"6",[282],[440],"tool output",[442],"claimed 60–90%; measured 0–90%",[444],"5 minutes",[446,448,450,451,453],[447],"7",[449],"context-mode",[440],[452],"claimed up to 98%; measured 20–98%",[454],"10 minutes",[456,458,463,464,466],[457],"8",[459,460],"Your own filter hooks, ",{"tag":190,"children":461},[462],"MAX_MCP_OUTPUT_TOKENS",[440],[465],"my build log: −97.5%",[467],"an hour",[469,471,472,474,476],[470],"9",[17],[473],"reads",[475],"claimed 98% in map mode",[454],[478,480,482,483,485],[479],"10",[481],"Serena",[473],[484],"mechanism; no neutral number",[486],"15 minutes",[488,490,492,493,494],[489],"11",[491],"Code intelligence plugins",[473],[405],[390],[496,497,499,500,502],[343],[498],"A lean CLAUDE.md, skills for the rest",[212],[501],"official: under 200 lines",[467],[504,506,508,509,510],[505],"13",[507],"Subagents for verbose work",[403],[405],[407],[512,514,516,517,519],[513],"14",[515],"token-savior",[473],[518],"measured −43%",[486],[521,523,525,526,528],[522],"15",[524],"Repo maps: Aider, code-review-graph",[473],[527],"claimed large; measured −5% on a small repo",[529],"varies",[531,533,537,538,540],[532],"16",[534],{"tag":190,"children":535},[536],"repomix --compress",[473],[539],"my run: −79% TS, +30% Vue",[390],[542,544,546,547,549],[543],"17",[545],"Context7",[473],[548],"mechanism; no number",[390],[551,553,555,556,558],[552],"18",[554],"claude-context",[473],[557],"claimed about 40%; measured 30–60% on monorepos",[559],"an hour, needs a vector DB",[561,563,565,566,568],[562],"19",[564],"Terse output rules: caveman, claude-token-efficient",[432],[567],"4–12% of output tokens",[390],[570,572,574,576,578],[571],"20",[573],"claude-code-router",[575],"price, not tokens",[577],"3–5x cheaper on routed turns",[467],{"type":165,"content":580},[581,582,585],"The only independent head-to-head test I found is by ",{"tag":239,"href":83,"children":583},[584],"ComputingForGeeks",", from April 2026: one repository (sindresorhus\u002Fky), Claude Code 2.1.116 and Sonnet 4.5, against a baseline of 284,473 tokens and $0.27. One repository is thin evidence, so I use it as a tiebreaker, not a verdict.",{"type":289,"head":587,"rows":594},[588,590,592],[589],"Tool",[591],"Change in total tokens",[593],"Note",[595,601,608,615,622,629,636,642],[596,597,599],[515],[598],"−43%",[600],"symbol index and memory over MCP",[602,604,606],[603],"claude-token-efficient",[605],"−40%",[607],"CLAUDE.md rules",[609,611,613],[610],"caveman",[612],"−38%",[614],"output style",[616,618,620],[617],"token-optimizer-mcp",[619],"−23%",[621],"MCP server",[623,625,627],[624],"alexgreensh\u002Ftoken-optimizer",[626],"−18%",[628],"PolyForm Noncommercial license",[630,632,634],[631],"code-review-graph",[633],"about −5%",[635],"small repo, the graph overhead eats the gain",[637,638,640],[282],[639],"0% on clean output, 60–90% on noisy logs",[641],"depends on the commands",[643,644,646],[449],[645],"−20% to −98%",[647],"depends on the workload",{"type":205,"level":206,"id":137,"text":138},{"type":205,"level":650,"text":651},3,"1. Measure first: \u002Fusage, \u002Fcontext and ccusage",{"type":165,"content":653},[654,656,657,659,660,662,663,666],{"tag":190,"children":655},[355]," shows the session’s tokens and, in current versions, a prompt cache line with the share of input served from cache, the number of misses and a likely cause for the last one. ",{"tag":190,"children":658},[383]," shows what fills the window right now: system prompt, tools, memory files and messages. For history across sessions, ",{"tag":239,"href":77,"children":661},[23]," (",{"tag":190,"children":664},[665],"npx ccusage@latest",", MIT) reads the local JSONL logs and prints daily, monthly, per-session and five-hour-block reports. Without a baseline you cannot tell a 40% tool from a placebo.",{"type":205,"level":650,"text":668},"2. \u002Fclear between tasks, \u002Fcompact with instructions",{"type":165,"content":670},[671,672,674,675,678,679,682,683,181],"The whole history is sent on every turn, so stale context from the last task is paid for again with every new message. ",{"tag":190,"children":673},[397]," starts fresh and costs nothing. ",{"tag":190,"children":676},[677],"\u002Fcompact Focus on the failing test and the diff"," keeps continuity, but it has to read the whole conversation to summarise it, so it is a large request of its own. Use ",{"tag":190,"children":680},[681],"\u002Frename"," before clearing if you want to come back later with ",{"tag":190,"children":684},[685],"\u002Fresume",{"type":205,"level":650,"text":687},"3. Keep the prefix stable so caching works",{"type":165,"content":689},[690],"Caching is the biggest discount on this list and it is on by default, which is why it is easy to break without noticing. The cache follows the order tools, system prompt, messages: change a tool definition and everything after it is written again, at the higher cache-write price. In practice that means not toggling MCP servers in the middle of a session and not editing CLAUDE.md halfway through a long task. The cache lifetime is an hour on a subscription and five minutes by default on an API key, so on the API a coffee break costs a full re-read of the context.",{"type":205,"level":650,"text":692},"4. Fewer MCP servers, deferred tools, CLIs where they exist",{"type":165,"content":694},[695,696,699,700,703,704,707,708,711,712,715,716,380,719,722,723,726],"Anthropic’s ",{"tag":239,"href":92,"children":697},[698],"tool search documentation"," gives a concrete number: five common servers (GitHub, Slack, Sentry, Grafana and Splunk) take about 55,000 tokens of definitions before any work starts, and deferred loading typically cuts that by more than 85%. It also notes that tool selection gets worse beyond 30 to 50 loaded tools. Claude Code defers MCP tools by default, but its ",{"tag":239,"href":89,"children":701},[702],"MCP docs"," list setups where tool search is off, including ",{"tag":190,"children":705},[706],"ENABLE_TOOL_SEARCH=false"," and a custom ",{"tag":190,"children":709},[710],"ANTHROPIC_BASE_URL",". Beyond that, disable servers you are not using in ",{"tag":190,"children":713},[714],"\u002Fmcp",", and prefer ",{"tag":190,"children":717},[718],"gh",{"tag":190,"children":720},[721],"aws"," or ",{"tag":190,"children":724},[725],"gcloud"," over an MCP wrapper for the same API, because a CLI adds no tool listing at all.",{"type":205,"level":650,"text":728},"5. Match effort and model to the task",{"type":165,"content":730},[731,732,735,736,739,740,743,744,748],"Thinking tokens are billed as output tokens, the most expensive kind. ",{"tag":190,"children":733},[734],"\u002Feffort"," lowers the reasoning budget on adaptive models and ",{"tag":190,"children":737},[738],"\u002Fmodel"," switches to a cheaper one; subagents can run on Haiku with ",{"tag":190,"children":741},[742],"model: haiku"," in their definition. The ",{"tag":172,"to":745,"children":746},"\u002Fblog\u002Fartificial-analysis-leaderboard-claude-opus-5-5",[747],"Opus 5.5 leaderboard numbers"," show the scale: at medium effort the model matched its predecessor’s max-effort score for about a quarter of the cost per task.",{"type":205,"level":206,"id":140,"text":141},{"type":205,"level":650,"text":751},"6. rtk",{"type":165,"content":753},[754,756,757,760,761,176,764,767,768,770,771,774],{"tag":239,"href":38,"children":755},[282]," is a Rust CLI that sits in front of shell commands through a PreToolUse hook (",{"tag":190,"children":758},[759],"rtk init -g",") and filters, groups, truncates and deduplicates the output of more than a hundred commands. The README reports about 70% for ",{"tag":190,"children":762},[763],"ls",{"tag":190,"children":765},[766],"tree",", about 80% for ",{"tag":190,"children":769},[192]," and 90% for ",{"tag":190,"children":772},[773],"cargo test",". The independent test found 0% on commands that were already quiet and 60–90% on noisy logs, which matches my build log. Two limits: the figures are output reduction, not bill reduction, and the hook only sees shell commands, so Claude Code’s built-in Read, Grep and Glob tools pass through untouched. Apache 2.0.",{"type":205,"level":650,"text":776},"7. context-mode",{"type":165,"content":778},[779,781],{"tag":239,"href":44,"children":780},[449]," takes a different route: tool output goes into a sandbox and a SQLite FTS5 index, and the agent searches it with BM25 instead of reading it whole. The project reports 315 KB of output shrinking to 5.4 KB, a Playwright snapshot going from 56 KB to 299 bytes and an access log from 45 KB to 155 bytes, and it keeps a session guide of at most 2 KB that survives compaction. The independent test measured 20–98% depending on the workload. It is licensed under the Elastic License 2.0, which is fine for your own use but is not an OSI open-source license.",{"type":205,"level":650,"text":783},"8. Your own filter hooks, and a cap on MCP output",{"type":165,"content":785},[786],"The official cost guide shows a PreToolUse hook that rewrites test commands so only failures reach the model. The same idea fits any command you run often. This is the version I would write for the build above:",{"type":190,"code":788},"#!\u002Fbin\u002Fbash\n# ~\u002F.claude\u002Fhooks\u002Fquiet-build.sh: a PreToolUse hook with \"matcher\": \"Bash\".\n# Rewrites the site build so the model sees summary and error lines, not 500 lines of routes.\ninput=$(cat)\ncmd=$(echo \"$input\" | jq -r '.tool_input.command')\n\nif [[ \"$cmd\" =~ ^npm\\ run\\ generate ]]; then\n  quiet=\"set -o pipefail; $cmd 2>&1 | perl -pe 's\u002F\\e\\[[0-9;]*m\u002F\u002Fg' | grep -vE 'ExperimentalWarning|trace-warnings|├─|└─|node_modules\u002F.cache'\"\n  echo \"$input\" | jq --arg c \"$quiet\" \\\n    '{hookSpecificOutput: {hookEventName: \"PreToolUse\", permissionDecision: \"allow\", updatedInput: (.tool_input + {command: $c})}}'\nelse\n  echo \"{}\"\nfi\n",{"type":165,"content":790},[791,792,795,796,799,800,803,804,807,808,811,812,814],"Run against the log from the table, it leaves exactly the 46 lines of the fourth row, and with ",{"tag":190,"children":793},[794],"pipefail"," a failing build still exits non-zero. Register it in settings.json under ",{"tag":190,"children":797},[798],"hooks.PreToolUse"," with ",{"tag":190,"children":801},[802],"\"matcher\": \"Bash\"",", as in the official example, and check it with ",{"tag":190,"children":805},[806],"\u002Fhooks",". Like that example it answers ",{"tag":190,"children":809},[810],"allow",", which also skips the permission prompt for the rewritten command, so keep the pattern narrow. For MCP servers the equivalent lever is ",{"tag":190,"children":813},[462],": Claude Code warns when a single tool result passes 10,000 tokens and allows 25,000 by default, and a lower cap stops one chatty server from flooding the window.",{"type":205,"level":206,"id":143,"text":144},{"type":205,"level":650,"text":817},"9. lean-ctx",{"type":165,"content":819},[820,822,823,826],{"tag":239,"href":41,"children":821},[17]," is a local Rust binary and MCP server that gives the agent ten ways to read a file, from the full text to a map of its structure or only its signatures, plus compression patterns for more than 95 shell commands. On the project’s own 50-file repository, counted with the GPT-4o tokenizer, 533,200 raw tokens became 8,000 in map mode and 14,000 in signatures mode, and re-reading an unchanged file from its cache costs about 13 tokens. It is the most ambitious tool here and the numbers are the author’s own, but the idea of reading structure first and bodies on demand is sound. ",{"tag":190,"children":824},[825],"lean-ctx wrap claude",", Apache 2.0.",{"type":205,"level":650,"text":828},"10. Serena",{"type":165,"content":830},[831,833,834,380,837,176,840,843,844,176,847,850],{"tag":239,"href":47,"children":832},[481]," wraps language servers for more than 40 languages in MCP tools such as ",{"tag":190,"children":835},[836],"find_symbol",{"tag":190,"children":838},[839],"find_referencing_symbols",{"tag":190,"children":841},[842],"replace_symbol_body",". Instead of grepping and reading three candidate files, the agent asks for one symbol and edits it in place. I found no neutral benchmark, but the mechanism is the same one Anthropic recommends in the next item, and it works in any MCP client. Install with ",{"tag":190,"children":845},[846],"uv tool install -p 3.13 serena-agent",{"tag":190,"children":848},[849],"serena init","; GPL-3.0.",{"type":205,"level":650,"text":852},"11. Code intelligence plugins",{"type":165,"content":854},[855,856,858],"Claude Code’s own code intelligence plugins bring the same idea without a third-party server: go to definition and find references through an installed language server. The ",{"tag":239,"href":86,"children":857},[245]," puts it plainly: one definition lookup replaces a grep followed by reading several candidate files, and the language server reports type errors after edits, which saves a compile round trip. For TypeScript, Python, Go or Rust projects this is the first read-side change I would make.",{"type":205,"level":206,"id":146,"text":147},{"type":205,"level":650,"text":861},"12. A lean CLAUDE.md, with skills for the rest",{"type":165,"content":863},[864,865,869],"CLAUDE.md is loaded into every session, so every line in it is paid for on every request, cached or not. The official advice is to keep it under 200 lines and move workflow-specific instructions, such as how to review a PR or run a migration, into ",{"tag":172,"to":866,"children":867},"\u002Fblog\u002Fcoding-agent-skills-workflow",[868],"skills",", which load only when they are used. A short skill that describes the architecture also saves the exploratory reads an agent does at the start of every task.",{"type":205,"level":650,"text":871},"13. Subagents for verbose work",{"type":165,"content":873},[874],"A subagent runs tests, reads logs or fetches documentation in its own context and returns a summary, so the verbose part never enters the main history. It still costs tokens, just not again on every later turn, and it can run on a small model. The opposite warning is in the same docs: agent teams use roughly seven times the tokens of a normal session when teammates run in plan mode, because each teammate keeps its own full context.",{"type":205,"level":206,"id":149,"text":150},{"type":205,"level":650,"text":877},"14. token-savior",{"type":165,"content":879},[880,882],{"tag":239,"href":50,"children":881},[515]," combines a symbol index over MCP, a memory store and compaction of bash output. It made the largest cut in the independent test, 43% of total tokens. The project’s own headline, 80% fewer active tokens across 96 tasks with Opus 4.7, is marked unverified by the author, and an earlier figure was withdrawn. I read that as a sign of honesty, not as a reason to trust the bigger number. MIT.",{"type":205,"level":650,"text":884},"15. Repo maps: Aider and code-review-graph",{"type":165,"content":886},[887,888,891,892,895,896,898],"A repo map gives the agent a ranked outline of the codebase instead of whole files. ",{"tag":239,"href":56,"children":889},[890],"Aider’s repo map"," builds it with tree-sitter and a graph ranking and fits it into a budget set by ",{"tag":190,"children":893},[894],"--map-tokens",", 1,000 tokens by default. ",{"tag":239,"href":53,"children":897},[631]," stores a call graph in SQLite and answers blast-radius questions about a change. It reports a median of about 63 times fewer tokens per question, but says itself that this compares a graph query with the whole corpus, which is an upper bound. In the independent test on a small repository it saved about 5%, because the graph overhead ate most of the gain. Worth it on large repositories, not on small ones. MIT.",{"type":205,"level":650,"text":900},"16. Repomix --compress",{"type":165,"content":902},[903,904,907,908,911,912,914],"Repomix packs a repository into one file for a model, with ",{"tag":190,"children":905},[906],"--token-count-tree"," to show where the tokens are, ",{"tag":190,"children":909},[910],"--remove-comments"," and an MCP mode. ",{"tag":190,"children":913},[257]," is useful for giving a model a signature-level overview of a TypeScript, Python or Go codebase in one go, as my run showed. Check the output on template-heavy formats such as Vue before you rely on it. MIT.",{"type":205,"level":650,"text":916},"17. Context7 for library docs",{"type":165,"content":918},[919,921,922,925,926,929],{"tag":239,"href":62,"children":920},[545]," fetches current, version-specific documentation for a library on request, either through ",{"tag":190,"children":923},[924],"ctx7"," CLI commands with a skill or through an MCP server (",{"tag":190,"children":927},[928],"npx ctx7 setup","). The saving is indirect and I have no number for it: a focused snippet instead of a web page, and fewer rounds of fixing an API the model remembered from an older version. MIT.",{"type":205,"level":650,"text":931},"18. claude-context",{"type":165,"content":933},[934,936],{"tag":239,"href":65,"children":935},[554]," indexes the codebase for hybrid search, BM25 plus vectors, so the agent can ask for the code that handles authentication and get the relevant chunks. It reports about 40% fewer tokens at the same retrieval quality, and the independent test saw 30–60% on monorepos. The cost is infrastructure: an embedding provider (OpenAI, VoyageAI, Gemini or a local Ollama) and a Milvus or Zilliz Cloud vector database. It pays off on large monorepos and is overkill below that. MIT.",{"type":205,"level":206,"id":152,"text":153},{"type":205,"level":650,"text":939},"19. Terse output rules: caveman and claude-token-efficient",{"type":165,"content":941},[942,943,945,946,948],"These tools change how the model writes, not what it reads. ",{"tag":239,"href":68,"children":944},[610]," is a skill with lite, full and ultra modes of clipped prose; ",{"tag":239,"href":71,"children":947},[603]," is an eight-rule CLAUDE.md. The careful numbers are small. For caveman, a JetBrains lab run over 86 tasks found 8.5% fewer output tokens with flat quality, the project’s own eval shows a 50% median on short questions and answers, and agentic sessions see high single digits, while its rule file adds about 1,000 input tokens. claude-token-efficient measured 4% fewer output tokens on Haiku, 12% on Sonnet and 7% on Opus. The −38% and −40% from the independent test are far above that, and I would not generalise from one repository. Cheap to try, and most useful when you pay for a lot of output.",{"type":205,"level":650,"text":950},"20. claude-code-router",{"type":165,"content":952},[953,955,956,958],{"tag":239,"href":74,"children":954},[573]," is a local gateway that sends Claude Code’s requests to other providers and models, such as DeepSeek, Gemini, Kimi or OpenRouter, by rule. It does not reduce tokens; it makes some of them cheaper, and the independent test reported three to five times lower cost on the turns it routed. It is last on this list for a reason: a router means a custom ",{"tag":190,"children":957},[710],", which is one of the setups where Claude Code turns MCP tool search off, and every switch to another model starts that model’s cache from zero. Measure the whole session, not just the routed turns. MIT.",{"type":205,"level":206,"id":155,"text":156},{"type":961,"ordered":962,"items":963},"list",false,[964,973,978,983],[965,968,969,972],{"tag":197,"children":966},[967],"LLMLingua and other prompt compressors."," ",{"tag":239,"href":80,"children":970},[971],"LLMLingua"," drops tokens that a small model judges unimportant and reports up to 20 times compression with little loss on prose, retrieval and reasoning prompts. For code an agent has to edit exactly, lossy compression is the wrong trade.",[974,977],{"tag":197,"children":975},[976],"Tools whose license does not fit."," alexgreensh\u002Ftoken-optimizer saved 18% in the independent test, but it is under PolyForm Noncommercial, which rules it out for client work.",[979,982],{"tag":197,"children":980},[981],"Anything without a before and after."," The independent test did not recommend nadimtuhin\u002Fclaude-token-optimizer and found the effect of claude-mem variable. Memory tools can save re-explaining, or they can inject stale notes into every session.",[984,987],{"tag":197,"children":985},[986],"Three tools on the same layer."," rtk, context-mode and lean-ctx all intercept shell output. Pick one per layer, or you end up debugging which hook rewrote what.",{"type":205,"level":206,"id":158,"text":159},{"type":961,"ordered":990,"items":991},true,[992,1001,1009,1011,1013,1015],[993,994,997,998,1000],"Run ",{"tag":190,"children":995},[996],"npx ccusage@latest daily"," and one normal session with ",{"tag":190,"children":999},[355]," at the end. Write the numbers down.",[1002,1003,1005,1006,1008],"Open ",{"tag":190,"children":1004},[383],", disable the MCP servers you did not use this week, and replace any that have a CLI (",{"tag":190,"children":1007},[718]," instead of a GitHub server).",[1010],"Cut CLAUDE.md to under 200 lines and move the workflows into skills.",[1012],"Install the code intelligence plugin for your main language, or Serena if you use several agents.",[1014],"Add one output filter: rtk if you want it ready-made, a 15-line hook like the one above if you want to see exactly what it drops.",[1016],"Repeat the same kind of session and compare. Keep what moved the number and uninstall the rest.",{"type":165,"content":1018},[1019,1020,1024],"For the reasoning behind this order, ",{"tag":172,"to":1021,"children":1022},"\u002Fblog\u002Fharness-engineering-coding-agents",[1023],"harness engineering"," covers how guides and sensors keep an agent’s context useful, not only small.",{"type":205,"level":206,"id":161,"text":162},{"type":961,"ordered":962,"items":1027},[1028,1031,1034,1037,1040,1043,1046,1049,1052,1055,1058,1061,1064,1067,1070,1073,1076,1079,1082,1085],[1029],{"tag":239,"href":38,"children":1030},[37],[1032],{"tag":239,"href":41,"children":1033},[40],[1035],{"tag":239,"href":44,"children":1036},[43],[1038],{"tag":239,"href":47,"children":1039},[46],[1041],{"tag":239,"href":50,"children":1042},[49],[1044],{"tag":239,"href":53,"children":1045},[52],[1047],{"tag":239,"href":56,"children":1048},[55],[1050],{"tag":239,"href":59,"children":1051},[58],[1053],{"tag":239,"href":62,"children":1054},[61],[1056],{"tag":239,"href":65,"children":1057},[64],[1059],{"tag":239,"href":68,"children":1060},[67],[1062],{"tag":239,"href":71,"children":1063},[70],[1065],{"tag":239,"href":74,"children":1066},[73],[1068],{"tag":239,"href":77,"children":1069},[76],[1071],{"tag":239,"href":80,"children":1072},[79],[1074],{"tag":239,"href":83,"children":1075},[82],[1077],{"tag":239,"href":86,"children":1078},[85],[1080],{"tag":239,"href":89,"children":1081},[88],[1083],{"tag":239,"href":92,"children":1084},[91],[1086],{"tag":239,"href":95,"children":1087},[94],[1089,1147,1198,1253],{"slug":1090,"published":1091,"minutes":1092,"category":7,"tags":1093,"keywords":1098,"about":1109,"sources":1116,"cover":1141,"og":1142,"expertise":98,"locales":1143,"lang":100,"title":1144,"description":1145,"coverAlt":1146},"agent-loop-explained","2026-09-27",10,[1094,1095,1096,9,1097],"Agent loop","Coding agents","Tool calling","Stop conditions",[1099,1100,1101,1102,1103,1104,1105,1106,1107,1108],"agent loop","agentic loop","AI agent loop","tool calling loop","Ralph loop coding agent","Claude Code \u002Floop","Claude Code \u002Fgoal","how to stop an AI agent from looping","ReAct agent pattern","stop_reason tool_use",[1110,1113,1114],{"name":1111,"url":1112},"Intelligent agent","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FIntelligent_agent",{"name":27,"url":28},{"name":9,"url":1115},"https:\u002F\u002Fcode.claude.com\u002Fdocs\u002Fen\u002Foverview",[1117,1120,1123,1126,1129,1132,1135,1138],{"title":1118,"url":1119},"Anthropic: Building effective agents (Dec 2024)","https:\u002F\u002Fwww.anthropic.com\u002Fengineering\u002Fbuilding-effective-agents",{"title":1121,"url":1122},"Yao et al.: ReAct: Synergizing Reasoning and Acting in Language Models (2022)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2210.03629",{"title":1124,"url":1125},"Claude API docs: Handling stop reasons","https:\u002F\u002Fplatform.claude.com\u002Fdocs\u002Fen\u002Fbuild-with-claude\u002Fhandling-stop-reasons",{"title":1127,"url":1128},"Anthropic: Effective harnesses for long-running agents (Nov 2025)","https:\u002F\u002Fwww.anthropic.com\u002Fengineering\u002Feffective-harnesses-for-long-running-agents",{"title":1130,"url":1131},"Simon Willison: Red\u002Fgreen TDD (Agentic Engineering Patterns)","https:\u002F\u002Fsimonwillison.net\u002Fguides\u002Fagentic-engineering-patterns\u002Fred-green-tdd\u002F",{"title":1133,"url":1134},"Geoffrey Huntley: Ralph Wiggum as a software engineer (Jul 2025)","https:\u002F\u002Fghuntley.com\u002Fralph\u002F",{"title":1136,"url":1137},"Claude Code docs: Run prompts on a schedule (\u002Floop)","https:\u002F\u002Fcode.claude.com\u002Fdocs\u002Fen\u002Fscheduled-tasks",{"title":1139,"url":1140},"Claude Code docs: Keep Claude working toward a goal (\u002Fgoal)","https:\u002F\u002Fcode.claude.com\u002Fdocs\u002Fen\u002Fgoal","\u002Fimages\u002Fblog\u002Fagent-loop-explained\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fagent-loop-explained\u002Fog.jpg",[100,101,102],"The agent loop, explained: how coding agents run, and how to make them stop","How the agent loop works in code, which stop conditions and budgets to enforce, and how outer loops like Ralph and Claude Code \u002Floop and \u002Fgoal behave.","A five-step cycle: context, model, tool call, result and a stop check that either ends the agent loop or feeds back into the context.",{"slug":1148,"published":1091,"minutes":1092,"category":7,"tags":1149,"keywords":1155,"about":1165,"sources":1173,"cover":1192,"og":1193,"expertise":98,"locales":1194,"lang":100,"title":1195,"description":1196,"coverAlt":1197},"mcp-tool-design-lessons-jira-server",[1150,1151,1152,1153,1154],"MCP","Tool design","Context engineering","Jira","Agents",[1156,1157,1158,1159,1160,1161,1162,1163,1164],"MCP tool design","MCP best practices","MCP tool descriptions","agent tool selection","MCP context bloat","how many tools should an MCP server have","MCP tool definition token cost","MCP error handling isError","Jira MCP server",[1166,1169,1172],{"name":1167,"url":1168},"Model Context Protocol","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FModel_Context_Protocol",{"name":1170,"url":1171},"Jira (software)","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FJira_(software)",{"name":1111,"url":1112},[1174,1177,1180,1183,1186,1189],{"title":1175,"url":1176},"Writing effective tools for agents – with agents (Anthropic)","https:\u002F\u002Fwww.anthropic.com\u002Fengineering\u002Fwriting-tools-for-agents",{"title":1178,"url":1179},"Introducing advanced tool use on the Claude Developer Platform (Anthropic)","https:\u002F\u002Fwww.anthropic.com\u002Fengineering\u002Fadvanced-tool-use",{"title":1181,"url":1182},"Code execution with MCP (Anthropic)","https:\u002F\u002Fwww.anthropic.com\u002Fengineering\u002Fcode-execution-with-mcp",{"title":1184,"url":1185},"MCP vs CLI: context window cost (Blocks.ai)","https:\u002F\u002Fblocks.ai\u002Fblog\u002Fmcp-vs-cli-context-window-cost",{"title":1187,"url":1188},"Demystifying evals for AI agents (Anthropic)","https:\u002F\u002Fwww.anthropic.com\u002Fengineering\u002Fdemystifying-evals-for-ai-agents",{"title":1190,"url":1191},"MCP 2026-07-28 specification: Tools","https:\u002F\u002Fmodelcontextprotocol.io\u002Fspecification\u002F2026-07-28\u002Fserver\u002Ftools","\u002Fimages\u002Fblog\u002Fmcp-tool-design-lessons-jira-server\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fmcp-tool-design-lessons-jira-server\u002Fog.jpg",[100,101,102],"MCP tool design: lessons from a 20-tool Jira server","MCP tool design that agents get right: token cost of tool definitions, when to merge tools, naming, concise output, errors that steer and a small selection eval.","Network diagram with a Jira MCP server at the hub and five satellites: search, create, transition, comments and test runs",{"slug":1199,"published":1091,"minutes":1200,"category":7,"tags":1201,"keywords":1206,"about":1215,"sources":1224,"cover":1247,"og":1248,"expertise":98,"locales":1249,"lang":100,"title":1250,"description":1251,"coverAlt":1252},"harness-engineering-coding-agents",8,[1202,1095,1203,1204,1205],"Harness engineering","Code quality","Mutation testing","TDD",[1023,1207,1208,1209,1210,1211,1212,1213,1214],"harness engineering coding agents","AI code quality","coding agent guardrails","mutation testing AI generated tests","red\u002Fgreen TDD with AI agents","guides and sensors coding agents","how to make AI agent pull requests mergeable","can I trust tests written by AI",[1216,1218,1221],{"name":1204,"url":1217},"https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FMutation_testing",{"name":1219,"url":1220},"Test-driven development","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FTest-driven_development",{"name":1222,"url":1223},"Static program analysis","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FStatic_program_analysis",[1225,1228,1231,1234,1237,1238,1241,1244],{"title":1226,"url":1227},"Birgitta Böckeler: Harness engineering for coding agent users (Apr 2026)","https:\u002F\u002Fmartinfowler.com\u002Farticles\u002Fharness-engineering.html",{"title":1229,"url":1230},"Birgitta Böckeler: Maintainability sensors for coding agents (May 2026)","https:\u002F\u002Fmartinfowler.com\u002Farticles\u002Fsensors-for-coding-agents.html",{"title":1232,"url":1233},"Simon Willison: Agentic Engineering Patterns","https:\u002F\u002Fsimonwillison.net\u002Fguides\u002Fagentic-engineering-patterns\u002F",{"title":1235,"url":1236},"Simon Willison: First run the tests","https:\u002F\u002Fsimonwillison.net\u002Fguides\u002Fagentic-engineering-patterns\u002Ffirst-run-the-tests\u002F",{"title":1127,"url":1128},{"title":1239,"url":1240},"Claude Code docs: How Claude remembers your project","https:\u002F\u002Fcode.claude.com\u002Fdocs\u002Fen\u002Fmemory",{"title":1242,"url":1243},"Stryker Mutator documentation","https:\u002F\u002Fstryker-mutator.io\u002Fdocs\u002F",{"title":1245,"url":1246},"Infection: command line options","https:\u002F\u002Finfection.github.io\u002Fguide\u002Fcommand-line-options.html","\u002Fimages\u002Fblog\u002Fharness-engineering-coding-agents\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fharness-engineering-coding-agents\u002Fog.jpg",[100,101,102],"Harness engineering: guides and sensors that make agent PRs mergeable","Harness engineering for coding agents: guides and sensors, where to run each check, red\u002Fgreen TDD, and mutation testing to verify the tests the agent wrote.","Concentric rings around a coding agent's model: behaviour, architecture fitness and maintainability harnesses, from outside in.",{"slug":1254,"published":1091,"minutes":1255,"category":7,"tags":1256,"keywords":1260,"about":1271,"sources":1281,"cover":1311,"og":1312,"expertise":98,"locales":1313,"lang":100,"title":1314,"description":1315,"coverAlt":1316},"mcp-2026-07-28-stateless-migration-guide",9,[1150,1257,1258,1259,1154],"Protocol migration","Stateless APIs","OAuth",[1261,1262,1263,1264,1265,1266,1267,1268,1269,1270],"MCP 2026-07-28","stateless MCP","MCP migration guide","MCP sessions removed","Mcp-Session-Id removed","MCP server\u002Fdiscover","MCP multi round-trip requests","how to migrate an MCP server to the new spec","MCP CIMD client ID metadata document","MCP roots sampling logging deprecated",[1272,1273,1276,1279],{"name":1167,"url":1168},{"name":1274,"url":1275},"Stateless protocol","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FStateless_protocol",{"name":1277,"url":1278},"Load balancing (computing)","https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FLoad_balancing_(computing)",{"name":1259,"url":1280},"https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FOAuth",[1282,1285,1288,1291,1294,1297,1299,1302,1305,1308],{"title":1283,"url":1284},"The 2026-07-28 Specification (MCP blog)","https:\u002F\u002Fblog.modelcontextprotocol.io\u002Fposts\u002F2026-07-28\u002F",{"title":1286,"url":1287},"MCP 2026-07-28 Key Changes","https:\u002F\u002Fmodelcontextprotocol.io\u002Fspecification\u002F2026-07-28\u002Fchangelog",{"title":1289,"url":1290},"MCP 2026-07-28: Versioning and Compatibility","https:\u002F\u002Fmodelcontextprotocol.io\u002Fspecification\u002F2026-07-28\u002Fbasic\u002Fversioning",{"title":1292,"url":1293},"MCP 2026-07-28: Streamable HTTP transport","https:\u002F\u002Fmodelcontextprotocol.io\u002Fspecification\u002F2026-07-28\u002Fbasic\u002Ftransports\u002Fstreamable-http",{"title":1295,"url":1296},"MCP 2026-07-28: Multi Round-Trip Requests","https:\u002F\u002Fmodelcontextprotocol.io\u002Fspecification\u002F2026-07-28\u002Fbasic\u002Fpatterns\u002Fmrtr",{"title":1298,"url":1191},"MCP 2026-07-28: Tools",{"title":1300,"url":1301},"MCP 2026-07-28: Client registration","https:\u002F\u002Fmodelcontextprotocol.io\u002Fspecification\u002F2026-07-28\u002Fbasic\u002Fauthorization\u002Fclient-registration",{"title":1303,"url":1304},"MCP 2026-07-28: Deprecated features","https:\u002F\u002Fmodelcontextprotocol.io\u002Fspecification\u002F2026-07-28\u002Fdeprecated",{"title":1306,"url":1307},"The 2026 MCP Roadmap","https:\u002F\u002Fblog.modelcontextprotocol.io\u002Fposts\u002F2026-mcp-roadmap\u002F",{"title":1309,"url":1310},"RFC 9207: OAuth 2.0 Authorization Server Issuer Identification","https:\u002F\u002Fwww.rfc-editor.org\u002Frfc\u002Frfc9207","\u002Fimages\u002Fblog\u002Fmcp-2026-07-28-stateless-migration-guide\u002Fcover.webp","\u002Fimages\u002Fblog\u002Fmcp-2026-07-28-stateless-migration-guide\u002Fog.jpg",[100,101,102],"MCP 2026-07-28 migration guide: what changes for stateless MCP servers","MCP 2026-07-28 removes sessions and the initialize handshake. What changes for server authors: _meta, server\u002Fdiscover, MRTR, auth and a migration checklist.","Pipeline of five migration steps for MCP 2026-07-28: upgrade the SDK, remove sessions, add server\u002Fdiscover, rewrite prompts as MRTR, test across instances",1790582553878]