Blog/AI agents

Top 20 ways to cut coding-agent tokens: rtk, lean-ctx, Serena and more, ranked by evidence

rtk, lean-ctx, context-mode, Serena and 16 more token savers for coding agents, ranked by evidence, with my own measurements on a real Nuxt codebase.

··17 min read

  • Claude Code
  • token usage
  • context engineering
  • developer tools
Bar chart falling from a 15,100-token build log to about 370 tokens after filtering, under the heading Top 20 token savers.

Key takeaways

  • Most token-saver percentages are output reduction on the commands a tool is best at, not bill reduction, so treat them as upper bounds.
  • The free built-in habits come first: measure with /usage and ccusage, /clear between tasks, keep the cached prefix stable, trim MCP servers and lower effort.
  • Shell output was the biggest leak I measured: a 15,100-token build log shrank to about 370 tokens, with every useful line kept, by dropping colour codes, warnings and route lists.
  • On the read side, symbol navigation (code intelligence plugins, Serena, lean-ctx) beats reading whole files. Repomix --compress cut TypeScript by 79% in my run but made Vue files 30% larger.
  • The one independent head-to-head test I found put token-savior, claude-token-efficient and caveman in front at 38–43%, on a single repository.

Most advice on cutting coding-agent costs comes down to a screenshot of a counter that says 90% saved. I wanted to know which of these tools hold up, so I went through the READMEs and benchmarks of more than twenty token savers, compared them with the one independent head-to-head test I could find, and measured two of the ideas on this site’s own codebase. This is the ranking that came out of it.

The ranking is written for Claude Code, because that is where I work and where the official documentation is most specific, but most tools on the list also plug into Cursor, Codex, OpenCode or any other agent that speaks MCP or runs shell hooks. For the background on why context size drives both cost and quality, start with context engineering for coding agents and prompt caching and model routing.

Where the tokens in an agent turn come from

Every request an agent sends carries four kinds of tokens, and each tool on this list works on one of them. The prefix is the system prompt, the tool definitions and CLAUDE.md. Reads are the files and search results the agent pulls in to understand the code. Tool output is what shell commands, test runners and MCP servers print back. Model output is the thinking and the answer. All four pile up in the history, which is sent again on every turn.

Where a turn’s tokens come fromFour columns. Prefix: tool definitions and CLAUDE.md, cut by tool search, fewer MCP servers, CLIs instead of MCP, a lean CLAUDE.md and a stable prefix for caching. Reads: files and search results, cut by code intelligence, Serena and lean-ctx, token-savior, repo maps and Repomix, and claude-context. Tool output: shell, tests and MCP, cut by rtk, context-mode, your own hooks, a cap on MCP output and subagents. Model output: thinking and answers, cut by effort, model choice, Haiku for subagents, caveman and terse style rules. Below them a bar: the history, re-sent every turn, kept small with clear, compact and subagents and kept cheap by a stable cached prefix.Where a turn’s tokens come fromand what cuts themPrefixtools, CLAUDE.mdtool searchfewer MCP serversCLI over MCPlean CLAUDE.mdstable for cachingReadsfiles, search resultscode intelligenceSerena, lean-ctxtoken-saviorrepo maps, Repomixclaude-contextTool outputshell, tests, MCPrtkcontext-modeyour own hooksMCP output capsubagentsModel outputthinking, answers/effort/modelHaiku for subagentscavemanterse style rulesHistory: all of the above is sent again on every turn/clear · /compact · subagents · a stable prefix keeps it cached
Each tool works on one of four token sources; the history multiplies all of them by the number of turns.

The prices are not symmetric, and that decides where savings matter. On Claude Opus 5.5 a cached input token costs $0.20 per million, a fresh one $4 and an output token $20, according to the prompt caching docs. A stable, cached prefix is almost free after the first request. What costs money is new content: fresh reads, fresh tool output and everything the model writes, thinking included. Claude Code’s own cost guide puts the average at about $13 per developer per active day, and below $30 for 90% of users, so the problem is rarely one big bill. It is a steady leak across many turns.

What I measured on this site

Two of the claims were cheap to test here. This site is a Nuxt 4 app with about 110 Vue, TypeScript and script files, and its build prints a lot. I ran Repomix 1.18.1 with its default o200k_base tokenizer on the source, once plain and once with --compress, which uses tree-sitter to keep signatures and drop function bodies. The README puts the reduction at about 70%.

Repomix 1.18.1 on this siteGrouped bars of token counts. TypeScript and script files, 62 files: 133,146 tokens plain, 28,402 with compress, 79% fewer. Vue components, 49 files: 72,298 plain, 94,096 with compress, 30% more. Whole source, 111 files: 205,115 plain, 122,132 with compress, 40% fewer.Repomix 1.18.1 on this sitetokens, o200k_baseTypeScript, scripts62 files133,14628,402 −79%Vue components49 files72,29894,096 +30%Whole source111 files205,115122,132 −40%plain pack--compress--compress, larger
Measured on this site’s source: compression works on TypeScript and backfires on Vue single-file components.

On the TypeScript and script files the claim holds and then some: 133,146 tokens became 28,402, a 79% cut. On the Vue single-file components it went the other way: 72,298 tokens became 94,096, 30% more, and 46 of the 49 files grew. Looking at the output, the compressed Vue files kept the full template and then repeated fragments of it as extra chunks between ⋮---- markers. Across the whole codebase the saving was 40%, not 70%.

The second test was the use case behind rtk and every other output filter on this list: a noisy command. A full npm run generate of this site writes 501 lines. I stripped the log in stages and estimated tokens as bytes divided by four, the same rough heuristic rtk uses.

StageBytesTokens (bytes/4)Lines
Raw log, as written to a file60,311~15,100501
Without ANSI colour codes44,426~11,100501
Without Node experimental warnings42,144~10,500481
Without the per-route and per-chunk lists1,479~37046
Only summary and error lines620~15512

Colour codes alone were a quarter of the log, even though it was redirected to a file. The lists of prerendered routes and built chunks were 96% of what remained, and on a green build none of it tells a model anything. Dropping those three things keeps every line a person would act on and removes 97.5% of the tokens. On a red build you need the error lines and a few lines around them, which is exactly what a good filter keeps.

The top 20, ranked

I ranked by four things, in this order: whether the saving is backed by something other than the author’s own benchmark, how much of a real session it touches, how long setup takes, and what it risks, from lossy output to a restrictive license. Built-in habits come first because they are free and documented; third-party tools follow, ordered by evidence.

#Tool or habitLayerEvidenceSetup
1/usage, /context, ccusageallthe measurement itselfminutes
2/clear and a focused /compacthistoryofficial docsnone
3A stable prefix for prompt cachingprefixofficial pricingnone
4Tool search, fewer MCP servers, CLIsprefixofficial: over 85% of tool definitionsminutes
5Effort and model choicemodel outputofficial docsnone
6rtktool outputclaimed 60–90%; measured 0–90%5 minutes
7context-modetool outputclaimed up to 98%; measured 20–98%10 minutes
8Your own filter hooks, MAX_MCP_OUTPUT_TOKENStool outputmy build log: −97.5%an hour
9lean-ctxreadsclaimed 98% in map mode10 minutes
10Serenareadsmechanism; no neutral number15 minutes
11Code intelligence pluginsreadsofficial docsminutes
12A lean CLAUDE.md, skills for the restprefixofficial: under 200 linesan hour
13Subagents for verbose workhistoryofficial docsnone
14token-saviorreadsmeasured −43%15 minutes
15Repo maps: Aider, code-review-graphreadsclaimed large; measured −5% on a small repovaries
16repomix --compressreadsmy run: −79% TS, +30% Vueminutes
17Context7readsmechanism; no numberminutes
18claude-contextreadsclaimed about 40%; measured 30–60% on monoreposan hour, needs a vector DB
19Terse output rules: caveman, claude-token-efficientmodel output4–12% of output tokensminutes
20claude-code-routerprice, not tokens3–5x cheaper on routed turnsan hour

The only independent head-to-head test I found is by ComputingForGeeks, from April 2026: one repository (sindresorhus/ky), Claude Code 2.1.116 and Sonnet 4.5, against a baseline of 284,473 tokens and $0.27. One repository is thin evidence, so I use it as a tiebreaker, not a verdict.

ToolChange in total tokensNote
token-savior−43%symbol index and memory over MCP
claude-token-efficient−40%CLAUDE.md rules
caveman−38%output style
token-optimizer-mcp−23%MCP server
alexgreensh/token-optimizer−18%PolyForm Noncommercial license
code-review-graphabout −5%small repo, the graph overhead eats the gain
rtk0% on clean output, 60–90% on noisy logsdepends on the commands
context-mode−20% to −98%depends on the workload

Ranks 1–5: built-in habits that cost nothing

1. Measure first: /usage, /context and ccusage

/usage shows the session’s tokens and, in current versions, a prompt cache line with the share of input served from cache, the number of misses and a likely cause for the last one. /context shows what fills the window right now: system prompt, tools, memory files and messages. For history across sessions, ccusage (npx ccusage@latest, MIT) reads the local JSONL logs and prints daily, monthly, per-session and five-hour-block reports. Without a baseline you cannot tell a 40% tool from a placebo.

2. /clear between tasks, /compact with instructions

The whole history is sent on every turn, so stale context from the last task is paid for again with every new message. /clear starts fresh and costs nothing. /compact Focus on the failing test and the diff keeps continuity, but it has to read the whole conversation to summarise it, so it is a large request of its own. Use /rename before clearing if you want to come back later with /resume.

3. Keep the prefix stable so caching works

Caching is the biggest discount on this list and it is on by default, which is why it is easy to break without noticing. The cache follows the order tools, system prompt, messages: change a tool definition and everything after it is written again, at the higher cache-write price. In practice that means not toggling MCP servers in the middle of a session and not editing CLAUDE.md halfway through a long task. The cache lifetime is an hour on a subscription and five minutes by default on an API key, so on the API a coffee break costs a full re-read of the context.

4. Fewer MCP servers, deferred tools, CLIs where they exist

Anthropic’s tool search documentation gives a concrete number: five common servers (GitHub, Slack, Sentry, Grafana and Splunk) take about 55,000 tokens of definitions before any work starts, and deferred loading typically cuts that by more than 85%. It also notes that tool selection gets worse beyond 30 to 50 loaded tools. Claude Code defers MCP tools by default, but its MCP docs list setups where tool search is off, including ENABLE_TOOL_SEARCH=false and a custom ANTHROPIC_BASE_URL. Beyond that, disable servers you are not using in /mcp, and prefer gh, aws or gcloud over an MCP wrapper for the same API, because a CLI adds no tool listing at all.

5. Match effort and model to the task

Thinking tokens are billed as output tokens, the most expensive kind. /effort lowers the reasoning budget on adaptive models and /model switches to a cheaper one; subagents can run on Haiku with model: haiku in their definition. The Opus 5.5 leaderboard numbers show the scale: at medium effort the model matched its predecessor’s max-effort score for about a quarter of the cost per task.

Ranks 6–8: filter tool output before the model reads it

6. rtk

rtk is a Rust CLI that sits in front of shell commands through a PreToolUse hook (rtk init -g) and filters, groups, truncates and deduplicates the output of more than a hundred commands. The README reports about 70% for ls and tree, about 80% for git diff and 90% for cargo test. The independent test found 0% on commands that were already quiet and 60–90% on noisy logs, which matches my build log. Two limits: the figures are output reduction, not bill reduction, and the hook only sees shell commands, so Claude Code’s built-in Read, Grep and Glob tools pass through untouched. Apache 2.0.

7. context-mode

context-mode takes a different route: tool output goes into a sandbox and a SQLite FTS5 index, and the agent searches it with BM25 instead of reading it whole. The project reports 315 KB of output shrinking to 5.4 KB, a Playwright snapshot going from 56 KB to 299 bytes and an access log from 45 KB to 155 bytes, and it keeps a session guide of at most 2 KB that survives compaction. The independent test measured 20–98% depending on the workload. It is licensed under the Elastic License 2.0, which is fine for your own use but is not an OSI open-source license.

8. Your own filter hooks, and a cap on MCP output

The official cost guide shows a PreToolUse hook that rewrites test commands so only failures reach the model. The same idea fits any command you run often. This is the version I would write for the build above:

#!/bin/bash
# ~/.claude/hooks/quiet-build.sh: a PreToolUse hook with "matcher": "Bash".
# Rewrites the site build so the model sees summary and error lines, not 500 lines of routes.
input=$(cat)
cmd=$(echo "$input" | jq -r '.tool_input.command')

if [[ "$cmd" =~ ^npm\ run\ generate ]]; then
  quiet="set -o pipefail; $cmd 2>&1 | perl -pe 's/\e\[[0-9;]*m//g' | grep -vE 'ExperimentalWarning|trace-warnings|├─|└─|node_modules/.cache'"
  echo "$input" | jq --arg c "$quiet" \
    '{hookSpecificOutput: {hookEventName: "PreToolUse", permissionDecision: "allow", updatedInput: (.tool_input + {command: $c})}}'
else
  echo "{}"
fi

Run against the log from the table, it leaves exactly the 46 lines of the fourth row, and with pipefail a failing build still exits non-zero. Register it in settings.json under hooks.PreToolUse with "matcher": "Bash", as in the official example, and check it with /hooks. Like that example it answers allow, which also skips the permission prompt for the rewritten command, so keep the pattern narrow. For MCP servers the equivalent lever is MAX_MCP_OUTPUT_TOKENS: Claude Code warns when a single tool result passes 10,000 tokens and allows 25,000 by default, and a lower cap stops one chatty server from flooding the window.

Ranks 9–11: read symbols, not whole files

9. lean-ctx

lean-ctx is a local Rust binary and MCP server that gives the agent ten ways to read a file, from the full text to a map of its structure or only its signatures, plus compression patterns for more than 95 shell commands. On the project’s own 50-file repository, counted with the GPT-4o tokenizer, 533,200 raw tokens became 8,000 in map mode and 14,000 in signatures mode, and re-reading an unchanged file from its cache costs about 13 tokens. It is the most ambitious tool here and the numbers are the author’s own, but the idea of reading structure first and bodies on demand is sound. lean-ctx wrap claude, Apache 2.0.

10. Serena

Serena wraps language servers for more than 40 languages in MCP tools such as find_symbol, find_referencing_symbols and replace_symbol_body. Instead of grepping and reading three candidate files, the agent asks for one symbol and edits it in place. I found no neutral benchmark, but the mechanism is the same one Anthropic recommends in the next item, and it works in any MCP client. Install with uv tool install -p 3.13 serena-agent and serena init; GPL-3.0.

11. Code intelligence plugins

Claude Code’s own code intelligence plugins bring the same idea without a third-party server: go to definition and find references through an installed language server. The cost guide puts it plainly: one definition lookup replaces a grep followed by reading several candidate files, and the language server reports type errors after edits, which saves a compile round trip. For TypeScript, Python, Go or Rust projects this is the first read-side change I would make.

Ranks 12–13: keep the main context small

12. A lean CLAUDE.md, with skills for the rest

CLAUDE.md is loaded into every session, so every line in it is paid for on every request, cached or not. The official advice is to keep it under 200 lines and move workflow-specific instructions, such as how to review a PR or run a migration, into skills, which load only when they are used. A short skill that describes the architecture also saves the exploratory reads an agent does at the start of every task.

13. Subagents for verbose work

A subagent runs tests, reads logs or fetches documentation in its own context and returns a summary, so the verbose part never enters the main history. It still costs tokens, just not again on every later turn, and it can run on a small model. The opposite warning is in the same docs: agent teams use roughly seven times the tokens of a normal session when teammates run in plan mode, because each teammate keeps its own full context.

Ranks 14–18: indexes, maps and documentation

14. token-savior

token-savior combines a symbol index over MCP, a memory store and compaction of bash output. It made the largest cut in the independent test, 43% of total tokens. The project’s own headline, 80% fewer active tokens across 96 tasks with Opus 4.7, is marked unverified by the author, and an earlier figure was withdrawn. I read that as a sign of honesty, not as a reason to trust the bigger number. MIT.

15. Repo maps: Aider and code-review-graph

A repo map gives the agent a ranked outline of the codebase instead of whole files. Aider’s repo map builds it with tree-sitter and a graph ranking and fits it into a budget set by --map-tokens, 1,000 tokens by default. code-review-graph stores a call graph in SQLite and answers blast-radius questions about a change. It reports a median of about 63 times fewer tokens per question, but says itself that this compares a graph query with the whole corpus, which is an upper bound. In the independent test on a small repository it saved about 5%, because the graph overhead ate most of the gain. Worth it on large repositories, not on small ones. MIT.

16. Repomix --compress

Repomix packs a repository into one file for a model, with --token-count-tree to show where the tokens are, --remove-comments and an MCP mode. --compress is useful for giving a model a signature-level overview of a TypeScript, Python or Go codebase in one go, as my run showed. Check the output on template-heavy formats such as Vue before you rely on it. MIT.

17. Context7 for library docs

Context7 fetches current, version-specific documentation for a library on request, either through ctx7 CLI commands with a skill or through an MCP server (npx ctx7 setup). The saving is indirect and I have no number for it: a focused snippet instead of a web page, and fewer rounds of fixing an API the model remembered from an older version. MIT.

18. claude-context

claude-context indexes the codebase for hybrid search, BM25 plus vectors, so the agent can ask for the code that handles authentication and get the relevant chunks. It reports about 40% fewer tokens at the same retrieval quality, and the independent test saw 30–60% on monorepos. The cost is infrastructure: an embedding provider (OpenAI, VoyageAI, Gemini or a local Ollama) and a Milvus or Zilliz Cloud vector database. It pays off on large monorepos and is overkill below that. MIT.

Ranks 19–20: output style and routing

19. Terse output rules: caveman and claude-token-efficient

These tools change how the model writes, not what it reads. caveman is a skill with lite, full and ultra modes of clipped prose; claude-token-efficient is an eight-rule CLAUDE.md. The careful numbers are small. For caveman, a JetBrains lab run over 86 tasks found 8.5% fewer output tokens with flat quality, the project’s own eval shows a 50% median on short questions and answers, and agentic sessions see high single digits, while its rule file adds about 1,000 input tokens. claude-token-efficient measured 4% fewer output tokens on Haiku, 12% on Sonnet and 7% on Opus. The −38% and −40% from the independent test are far above that, and I would not generalise from one repository. Cheap to try, and most useful when you pay for a lot of output.

20. claude-code-router

claude-code-router is a local gateway that sends Claude Code’s requests to other providers and models, such as DeepSeek, Gemini, Kimi or OpenRouter, by rule. It does not reduce tokens; it makes some of them cheaper, and the independent test reported three to five times lower cost on the turns it routed. It is last on this list for a reason: a router means a custom ANTHROPIC_BASE_URL, which is one of the setups where Claude Code turns MCP tool search off, and every switch to another model starts that model’s cache from zero. Measure the whole session, not just the routed turns. MIT.

What I would skip, or use with care

  • LLMLingua and other prompt compressors. LLMLingua drops tokens that a small model judges unimportant and reports up to 20 times compression with little loss on prose, retrieval and reasoning prompts. For code an agent has to edit exactly, lossy compression is the wrong trade.
  • Tools whose license does not fit. alexgreensh/token-optimizer saved 18% in the independent test, but it is under PolyForm Noncommercial, which rules it out for client work.
  • Anything without a before and after. The independent test did not recommend nadimtuhin/claude-token-optimizer and found the effect of claude-mem variable. Memory tools can save re-explaining, or they can inject stale notes into every session.
  • Three tools on the same layer. rtk, context-mode and lean-ctx all intercept shell output. Pick one per layer, or you end up debugging which hook rewrote what.

A starter stack for one afternoon

  1. Run npx ccusage@latest daily and one normal session with /usage at the end. Write the numbers down.
  2. Open /context, disable the MCP servers you did not use this week, and replace any that have a CLI (gh instead of a GitHub server).
  3. Cut CLAUDE.md to under 200 lines and move the workflows into skills.
  4. Install the code intelligence plugin for your main language, or Serena if you use several agents.
  5. Add one output filter: rtk if you want it ready-made, a 15-line hook like the one above if you want to see exactly what it drops.
  6. Repeat the same kind of session and compare. Keep what moved the number and uninstall the rest.

For the reasoning behind this order, harness engineering covers how guides and sensors keep an agent’s context useful, not only small.

Sources

Frequently asked questions

What is the best tool to reduce Claude Code token usage?

There is no single one, because each tool works on a different part of a turn. Start with the free built-ins: /clear between tasks, a stable prefix for caching, fewer MCP servers and lower effort. Then add one output filter such as rtk or context-mode and one read-side tool such as a code intelligence plugin or Serena, and measure before and after.

Does rtk really save 90% of tokens?

On noisy commands such as test runs and build logs it removes 60–90% of the output, and an independent test confirmed that range. On commands whose output is already short it saves close to nothing, and it does not touch Claude Code’s built-in Read, Grep and Glob tools, so the effect on a whole session is much smaller than the headline.

Is it safe to compress context for a coding agent?

Methods that keep structure are safe: filtering noise out of logs, reading signatures before bodies and looking up symbols through a language server. Lossy prompt compression that drops individual tokens, such as LLMLingua, suits prose and retrieval, but not code the agent has to edit exactly.

How do I measure my token usage in Claude Code?

Use /usage for the current session, including its prompt cache hit rate, and /context to see what fills the context window. For history across sessions, ccusage reads the local logs and reports usage by day, month, session or five-hour block.

Sounds like what you need?

Tell me about your project or role – I’d love to hear from you.