Tools/AI agents

Aider review: git-first pair programming in the terminal

A review of Aider 0.86.2, an Apache-2.0 terminal pair programmer whose benchmark ranks models honestly and whose release cadence has stopped.

Type
Coding agent
Pricing
Free · BYO API key

··10 min read

  • Coding agent
  • Terminal
  • Git workflow
  • BYO key
  • Open source
Cover art for the Aider review: a terminal session turning a single request into a row of git commits

Key takeaways

  • Aider 0.86.2 was published on 12 February 2026 and is the only release since August 2025, with one maintainer listed on PyPI.
  • Its polyglot benchmark runs 225 Exercism exercises and prints the dollar cost of the full run next to the score, which is more than most model leaderboards disclose.
  • The newest benchmark entries are dated 3 October 2025, so the leaderboard cannot rank the models shipping in 2026.
  • Every successful edit becomes a git commit with a generated message, which keeps review ordinary and the history noisy until it is squashed.
  • Context is assembled by the engineer rather than the tool, so token spend stays predictable but cross-file archaeology stays manual.

Aider is an Apache-2.0 terminal pair programmer: the engineer describes the change, Aider edits the files in front of it and commits the result. The position taken here is that it is still the best-engineered of the free coding CLIs for one specific job — controlled, reviewable edits inside an existing repository — and that its problem in 2026 is momentum rather than quality: one release since August 2025 and a benchmark page whose newest entries are a year old.

It competes with Claude Code, Cursor, Cline and the editor-integrated assistants, but on different terms. Those products wrap a model in an agent that plans and runs tasks; Aider offers an edit protocol. The engineer keeps the model choice, the editor and the shape of the history, which is exactly why it appeals to teams with an existing review process and why it disappoints anyone expecting the tool to finish the work alone.

What it is

The design is narrow and deliberate. Aider runs in a shell inside a git repository, builds a map of the codebase to decide what the model needs to see, and applies the model's answer as a file edit rather than as chat prose. The conversation is the input; the diff is the product.

  • Apache-2.0 licensed, shipped as aider-chat on PyPI, current release 0.86.2 of 12 February 2026, Python 3.10 to 3.12.
  • About 49,400 stars, 5,000 forks and 13,138 commits on GitHub, with roughly 1,400 open issues and 517 open pull requests.
  • Providers are chosen per session or per config file: OpenAI, Anthropic, Gemini, DeepSeek, xAI, Azure, Bedrock, Vertex, OpenRouter, GitHub Copilot, Ollama or any OpenAI-compatible endpoint.
  • The repository map is built from tree-sitter parsing and travels with every request, so the model sees definitions without the whole tree being added to the chat.
  • Edits are applied in the edit format the model was chosen for: whole file, diff, fenced diff, or architect mode, which splits planning and editing across two models.
  • Every successful change becomes a git commit with a generated message, and lint and test commands can run after each edit with failures fed back into the chat.
  • Watch mode reacts to AI! and AI? comments in the files, so a change can be requested from inside the editor without switching windows.

Two things follow from that design. Review stays ordinary git: the unit of AI work is a commit, so revert, bisect and blame keep working on machine-written code. And Aider holds no state beyond the repository — no cloud project, no server-side index, no account — which makes adoption cheap and leaves context management entirely with the operator.

How it works

Every turn sends three things to the model: the chat history, the files added to the session, and the repository map. The map is a ranked outline of definitions across the repository, sized to a token budget that the startup banner prints; it is how Aider works on a codebase larger than the context window without pretending to have read all of it.

How one Aider turn runsEvery turn sends the request, the repository map and the files added to the session to the model, which answers in an edit format; Aider applies the diff, runs the lint and test commands and commits the change. Failures from lint or tests go back into the chat and the loop repeats. The engineer chooses which files are added to the session.One Aider turnthe diff is the productEVERY TURNRequestplain wordsContextmap plus added filesModeldiff or architectApplythe edit landsAFTER THE EDITfailures re-enter the chatLint and testnon-zero exits loop backGit commitgenerated messageOrdinary gitdiff, revert, bisect, blameOnly the files the engineer adds are resent; the repository map is rebuilt each turn.Aider reports token limits, it never enforces them
The model never sees the whole repository: it sees a map, the files in the session, and the diff it produced.

The loop closes with the checks the engineer already trusts. Lint and test commands run after each edit, a non-zero exit comes back into the chat as an error to fix, and the commit follows the fix. Aider enforces nothing: its own troubleshooting page states that it only reports token-limit errors returned by the provider and that the token counts it prints are estimates.

Getting started

Installation is two commands. The session below is the whole workflow — start in a repository, name the model, hand over the files to edit, and attach the lint and test commands that would be run anyway. Keys come from the environment or from the api-key flag.

python -m pip install aider-install
aider-install

cd your-repo
export ANTHROPIC_API_KEY=sk-ant-...

# name the files to edit; everything else stays out of the context
aider src/billing/invoice.py src/billing/tax.py \
  --model sonnet \
  --lint-cmd "ruff check" \
  --test-cmd "pytest -q" --auto-test

Two habits separate a cheap session from an expensive one. Add only the files that will actually be edited, because added files are resent on every turn; and prefer models that answer in a diff over models asked to rewrite whole files, because output tokens are what a large change runs into first. When a change spans more files than one prompt should hold, architect mode puts a planning model and an editor model on the job.

What it actually costs

The tool is free, so the entire bill lands on the API account and model choice becomes the only cost control that matters. The project's own benchmark prints the cost of a full run next to the score, which makes it the least hypothetical cost data available for this kind of work:

  • A full 225-exercise run with gpt-5 at high reasoning effort cost $29.08 and passed 88.0%.
  • The same run with Gemini 2.5 Pro and 32,000 thinking tokens cost $49.88 for 83.1%.
  • DeepSeek-V3.2 reached 74.2% in reasoner mode for $1.30, and 70.2% in chat mode for $0.88.
  • claude-3-7-sonnet without thinking passed 60.4% for $17.72, a fair illustration that spending more is not the same as editing better.
  • Aider's homepage reports roughly 15 billion tokens a week processed by its users; the number is self-reported, but it sets the scale of real usage.

A working day sits far below those figures, since a benchmark re-runs 225 exercises end to end, but the shape holds: every turn pays for the repository map plus the files in the session, whether or not the provider caches the prefix. Session hygiene is the budget.

The benchmark behind the advice

The published leaderboard is unusual in this field because it measures the model rather than the tool: 225 Exercism exercises across C++, Go, Java, JavaScript, Python and Rust, solved end to end, with two pass rates, the share of responses in a well-formed edit format, and the dollar cost of the run. The methodology is stated down to the commit hash, and it is why the project's model recommendations carry weight.

ModelPass rateCost of the full runEdit format
gpt-5 at high reasoning88.0%$29.08diff
Gemini 2.5 Pro, 32k thinking83.1%$49.88diff-fenced
o3 high with a gpt-4.1 editor78.2%$17.55architect
DeepSeek-V3.2 reasoner74.2%$1.30diff
claude-3-7-sonnet, no thinking60.4%$17.72diff

The caveat is freshness. The newest entries on that page are dated 3 October 2025, so a reader in October 2026 is looking at a year-old ranking of models that have since been replaced. The method remains valuable; the table does not, and anyone picking a model from it is shopping on last year's shelf.

Where it falls short

The weaknesses are structural. Context is assembled by hand: the repository map surfaces definitions, but a change crossing an unfamiliar subsystem still depends on the engineer knowing what to add, and nothing in the tool plans work across files. The feature set stops at editing — the documentation index covers linting, tests, voice, images and watch mode, and has no page for MCP servers, plugins or a skill system, so anything beyond file edits is wired up by the operator. Then there is pace: PyPI lists a single maintainer for the package, the previous release took six months, and about 1,400 issues and 517 pull requests are open.

ToolWhere it runsModel choiceWhat you pay
AiderTerminal, any editorAny provider or local modelFree, API usage only
Claude CodeTerminalClaude modelsAPI usage or subscription
CursorIts own editorVendor-hosted modelsSubscription
ClineVS Code extensionMost providers, BYO keyFree, API usage only

Read that table as a claim about control rather than capability. Claude Code and Cursor will do more of a task unprompted, and the price is a narrower model menu and, in Cursor's case, an editor owned by the vendor. Aider's offer is the mirror image — maximum control of the loop, minimum product around it — and a team that already reviews every diff is exactly the team that wants it.

Verdict

Aider is the right terminal agent for engineers who want an edit protocol instead of an autonomous colleague, and the wrong one for anyone who expects a tool to carry a task from issue to merged pull request. Its engineering — edit formats, repository map, benchmark, git discipline — remains ahead of most of the field. Its project dynamics are the risk, and the licence does not mitigate them.

  1. Choose it when every change must be a reviewable commit and AI edits should look exactly like human edits in the log.
  2. Choose it when model choice is a requirement: regulated data, an existing enterprise agreement, or a local model behind Ollama.
  3. Choose it when the engineer is willing to decide what belongs in context, because that responsibility never moves into the tool.
  4. Do not choose it if the task needs an agent that plans multi-step work, drives a browser, or extends through plugins and MCP.
  5. Do not choose a model from the published leaderboard without rerunning it; its newest entries are from October 2025.

One more consideration is reversibility. Nothing in the workflow is proprietary — files, prompts, configuration and history all live in the repository — so adopting the tool costs nothing to undo. That is a stronger argument than any benchmark, and it does not hold for an agent that lives inside an editor. For teams wiring Aider into a larger harness, the surrounding patterns are covered in the notes on harness engineering for coding agents.

Aider never enforces token limits, it only reports token limit errors from the API provider. The token counts that aider reports are estimates.

Sources

  1. Aider website
  2. Aider LLM leaderboards
  3. Aider linting and testing
  4. Aider token limits
  5. Aider repository on GitHub
  6. aider-chat on PyPI

Frequently asked questions

Is Aider free to use?

The tool itself is free: Apache-2.0, no account, no seat, no feature gate. You bring an API key and pay the provider. The benchmark's own cost column puts a full 225-exercise run at $0.88 with DeepSeek chat and $29.08 with gpt-5 at high reasoning, which is the most concrete pricing data available for this kind of work.

Aider or Claude Code?

Aider wins when model choice, an existing editor and commit-per-change review are requirements; it gives up task planning, plugins and MCP. Claude Code does more of a task unprompted and costs either API tokens or a subscription, with a narrower model menu. Teams that already review every diff tend to prefer the first, teams that want the tool to own the whole task tend to prefer the second.

How does Aider decide what the model sees?

Three inputs go out every turn: the chat history, the files explicitly added to the session, and a repository map built from tree-sitter parsing of the codebase. The map is sized to a token budget, which the startup banner prints. Because added files are resent each turn, the practical controls are /tokens, /drop and /clear.

Does Aider work with local models?

Yes, through Ollama, LM Studio or any OpenAI-compatible endpoint, and local models are the one way to run it at zero API cost. The trade is edit quality: weaker models make malformed edits more often, and Aider supports whole-file, diff and fenced-diff edit formats so a weaker model can be given the easier job. The project's leaderboard shows large quality gaps between model tiers on the same exercises.

Sounds like what you need?

Tell me about your project or role – I’d love to hear from you.