Tools/AI agents
Claude Code: the terminal coding agent, reviewed
An engineering review of Claude Code: the extension surface, the real cost per developer, and the exact boundary of the Bash sandbox.
- Type
- Coding agent
- Pricing
- Free · pay per API token
Balázs Csorba··10 min read
- Terminal agent
- Hooks
- Subagents
- MCP
- Sandbox

Key takeaways
- Version 2.1.292 shipped on 6 October 2026 and the changelog lists releases almost daily, so pinning a version for CI is standing work rather than a one-off.
- Six mechanisms extend the loop — CLAUDE.md, skills, subagents, MCP, hooks and plugins — and using the wrong one is the most common configuration mistake.
- Permission rules are enforced by the client, not the model, which makes a PreToolUse hook the only guard that still holds in bypassPermissions.
- The Bash sandbox covers Bash, PowerShell and Monitor commands only; file tools, MCP servers, hooks and language servers all run outside it.
- Anthropic's own cost documentation puts average enterprise spend at about $13 per developer per active day, with 90% of users below $30.
Claude Code is Anthropic's terminal-first coding agent: a CLI that reads a repository, edits files, runs commands and iterates on the result, with the Claude models behind it. On the evidence of the documentation it is the most complete agent harness on the market, and the reason is not the model but the six extension mechanisms wrapped around the loop. The cost of that completeness is a configuration surface large enough to get wrong, and a release cadence — 2.1.292 landed on 6 October 2026 — that turns version pinning into standing maintenance. Worth adopting, provided the security section below is read before the first unattended run.
It competes with Cursor's editor-native agent, GitHub Copilot's agent mode and the Codex CLI, and it is the only one of the four where the agent is the application rather than a feature of an editor. That matters on a large repository, where the agent gets the whole working tree, the shell and the git history rather than whatever happens to be open in a tab. The trade is the interface: a terminal is a poor place to read a diff, which is why the VS Code and JetBrains extensions exist at all.
What it actually is
The agentic loop is the product. A task moves through gather context, take action, verify results, over and over, with the model choosing the next tool call and the harness supplying the tools: file operations, search, shell execution, web search, and code intelligence through language-server plugins. Everything below the loop — settings files, skills, subagents, hooks — exists to make that loop cheaper, safer or broader.
- Version 2.1.292, published 6 October 2026; the changelog lists releases up to that point almost daily.
- Runs in the terminal, as VS Code and JetBrains extensions, on the desktop, in the browser, and on Amazon Bedrock and Google Cloud.
- Built-in tools cover file operations, search, shell, web search and code intelligence; the terminal CLI, VS Code and JetBrains also accept third-party providers.
- Permission modes are
default(labelled manual),acceptEdits,plan,auto,dontAskandbypassPermissions. - Auto mode is the built-in starting mode for interactive terminal and VS Code sessions on Pro, Max and Team: a separate classifier model reviews each action instead of a human.
- The Bash sandbox uses operating-system primitives — Seatbelt on macOS, bubblewrap on Linux — and applies to Bash, PowerShell and Monitor commands only.
- Bundled skills ship in the box, including
/doctor,/code-review,/batch,/debugand/loop.
How it works: one context window
Everything the model knows about a session lives in a single context window: CLAUDE.md and any AGENTS.md, auto memory, MCP tool names, skill descriptions, every file read, every tool result and the transcript itself. Path-scoped rules load when their trigger file is read. When the window fills, the session compacts — the conversation is replaced with a structured summary, the root CLAUDE.md and unscoped rules are re-injected from disk, up to five recently modified files are re-read, and invoked skill bodies come back capped at 5,000 tokens each and 25,000 in total.
The consequence for a bill is that re-sent history dominates, not the answer. A real session screen reports hundreds of thousands of cached input tokens against a few thousand output tokens, and cache reads are billed at a tenth of the input rate on current models. Anyone trying to reduce spend is really deciding what enters the window, which is why the documentation pushes MCP Tool Search, subagents and skills over simply asking the model to be brief.
The extension surface
Six mechanisms extend the loop, and the documentation is unusually clear that they are not interchangeable. CLAUDE.md is always-on context. A skill is knowledge or a workflow loaded on demand, following the Agent Skills open standard. A subagent is an isolated context that returns a summary instead of its transcript. MCP connects external services. A hook is a shell command, HTTP request, MCP tool call, single-turn prompt or agent that fires on a lifecycle event. Plugins bundle the lot for distribution across a team.
| Mechanism | What it is | Context cost | Deterministic? |
|---|---|---|---|
| CLAUDE.md | Always-on project instructions | Injected at session start | No, the model decides to follow it |
| Skill | Markdown knowledge or workflow | Description at start, body on use | No |
| Subagent | Isolated loop returning a summary | Only the summary returns | No |
| MCP server | External tools and data | Names at start, schemas on use | No |
| Hook | Script or model call on an event | Zero unless it returns output | Yes, the event always fires |
| Plugin | Bundle of skills, hooks, agents, MCP | Whatever the bundle contains | Depends on contents |
One sentence from the permissions documentation decides the whole design: permission rules are enforced by Claude Code, not by the model. An instruction like never edit .env in CLAUDE.md is a request; a PreToolUse hook that denies the edit is enforcement. Any team treating the first as a control has misread the threat model.
Getting started properly
Installation is a shell script and a login; there is no project to set up. The part of a first session worth getting right is the settings file, because that is where a team converts prompt instructions into enforced rules. This project-level configuration allowlists the commands it trusts, blocks a git push, formats every edit and switches on the Bash sandbox:
{
"permissions": {
"allow": [
"Bash(npm run *)",
"Bash(git commit *)",
"Read",
"Edit(src/**)"
],
"deny": [
"Bash(git push *)",
"Read(.env)"
]
},
"sandbox": {
"enabled": true,
"allowWrite": ["src"],
"allowedDomains": ["registry.npmjs.org"]
},
"hooks": {
"PostToolUse": [
{
"matcher": "Edit|Write",
"hooks": [
{ "type": "command", "command": "jq -r '.tool_input.file_path' | xargs npx prettier --write" }
]
}
]
}
}
Two details in that file are easy to get wrong. Bash(git commit *) matches git commit -m 'x' but not git -C . commit -m 'x', because the wildcard stands in for whatever text sits in its place. And a hook that exits 0 with no output has not approved anything — it has declined to decide, and the call continues through the normal permission flow. What does hold everywhere is a deny returned by a PreToolUse hook: it still blocks the tool in bypassPermissions mode, which is the one guarantee that survives a developer who has switched their own prompts off.
What it costs
The client is free to install and Claude Code is not part of the Free plan. Model access comes from a Claude subscription, a Team or Enterprise seat, or the Anthropic API billed per token. What makes the bill hard to predict is that subscriptions meter against rolling usage windows rather than publishing a token allowance.
| Route | Price | Claude Code | Shape |
|---|---|---|---|
| Free | $0 | Not included | Chat only, no terminal agent |
| Pro | $20 a month, $17 billed annually | Yes | Rolling session and weekly limits |
| Max | From $100 a month | Yes | 5x or 20x Pro usage |
| Team standard seat | $20 a seat annually, $25 monthly | Yes | Mix and match with premium seats |
| Team premium seat | $100 a seat annually, $125 monthly | Yes | Five times the standard seat usage |
| Enterprise | $20 a seat annually plus usage at API rates | Yes | SSO, SCIM, audit logs, spend limits |
| Anthropic API | Per token, no minimum | Yes | Hard ceiling, billed directly |
On the per-token route the current list prices are $4 per million input tokens and $20 per million output tokens for Opus 5.5, $2 and $10 for Sonnet 5.5, and $1 and $5 for Haiku 4.5, with cache reads at $0.20, $0.20 and $0.10. The cheapest lever is therefore the model: the same task on Sonnet 5.5 instead of Opus 5.5 halves the bill at list price, and Haiku 4.5 is a quarter of it. The documentation pushes the same point from the other end — moving long instructions out of CLAUDE.md and into skills, keeping MCP servers few, and pushing verbose work into subagents so their transcripts never enter the main window.
Security and isolation
The security model is the strongest part of the design and the easiest to misuse. Rules are matched by Claude Code rather than by the model, precedence puts a deny at any scope above an allow at any other, and managed settings outrank user settings, project settings and command-line flags. Deny rules hold in bypassPermissions, and a removal such as rm -rf / is refused even when an allow rule or a hook permits it. That is the behaviour you want from a circuit breaker.
The sandbox is the second layer, and its scope is narrower than the name suggests:
- Bash, PowerShell and Monitor commands and their child processes, on macOS, Linux and WSL2 — native Windows runs unsandboxed.
- File tools such as Read, Edit, Write and WebFetch, which follow permission rules instead; a sandbox
denyReadentry does not stop Read. - MCP servers, command hooks, plugin monitors, language servers and helper commands, all of which run with the session's full access.
- Commands that fail inside the sandbox can be retried unsandboxed through
dangerouslyDisableSandbox; settingallowUnsandboxedCommandsto false removes the escape hatch. - Trust is per directory and lasts one session: a project-level subagent's frontmatter hooks do not run until the workspace trust dialog is accepted, and a
-prun does not count as accepting it.
Where it weakens
The honest weaknesses come before the comparison. Claude Code is fast-moving and configuration-heavy, the model behind it is not yours to tune, and the agent has shell access by design.
- A release a day means behaviour can change under a CI job.
--bareexists to strip hooks, skills, commands, subagents, plugins, MCP servers, auto memory and CLAUDE.md for reproducible scripted runs, and it is the right default there. - Context is the scarce resource. Compaction drops the middle of the conversation, only five files are re-read afterwards, and skill bodies are truncated from the start — so the important instruction at the bottom of a long SKILL.md is the one that disappears.
- Descriptions of model-invocable skills load on every request, so vague or overlapping descriptions make the model load the wrong skill or miss the useful one.
- Subscriptions meter in time windows, not dollars. One heavy morning can exhaust the session limit with a week of allowance still unused, and only API billing offers a hard ceiling.
- It is closed. Anthropic's models are the product; only the terminal CLI, VS Code and JetBrains accept a third-party provider, and agent features are not portable to another harness.
| Option | What it is | Entry price | Main trade-off |
|---|---|---|---|
| Claude Code | Terminal agent, the application itself | $20 a month on Pro | Highest ceiling for repository-scale work, largest configuration surface |
| Cursor | Editor fork with an agent mode | $20 a month on Pro | Better diff and inline ergonomics, but the agent lives inside the editor |
| GitHub Copilot | IDE extension, completions plus agent | $10 a month on Pro | Cheapest entry point, least autonomy of the four |
| Codex CLI | Terminal agent on OpenAI models | Bundled with ChatGPT plans | A different model family, without the Claude-specific harness features |
The short version: for long-running, repository-scale work the terminal harness wins on capability; for line-by-line work an editor-native agent on a cheaper seat is the better buy. Running both is a defensible $30 a month, and most teams who end up there describe it as in-editor work for the small changes and the terminal for everything that spans a repository.
Verdict
Claude Code is worth adopting on one condition: the team writes the guardrails down and puts them in version control. Without a committed settings file, a PreToolUse hook on the protected paths and an explicit decision about unattended runs, the tool is faster than a reviewer and less careful than one.
- Adopt it when the work is repository-scale: migrations, refactors, multi-file changes that end in a verification step.
- Adopt it when the budget can carry one seat for a heavy user and one for a light user — the windows, not the seats, are the binding constraint.
- Adopt the permission and hook model seriously. It is the only difference between an agent and a supervised agent.
- Do not make it the only tool. Keep an editor-native completion product if most of the day is writing lines rather than changing systems.
- Do not run it unattended on a machine you care about without a container, and never with
--dangerously-skip-permissionsoutside one.
Permission rules are enforced by Claude Code, not by the model. Instructions in your prompt or CLAUDE.md shape what Claude tries to do, but they do not change what Claude Code allows.
Sources
- Claude Code documentation: overview
- Claude Code documentation: how Claude Code works
- Claude Code documentation: extend Claude Code
- Claude Code documentation: hooks reference
- Claude Code documentation: configure permissions
- Claude Code documentation: sandboxing
- Claude Code documentation: manage costs effectively
- Claude Code documentation: explore the context window
- Claude Code changelog
- Claude Platform pricing: model list prices and prompt caching
- Anthropic pricing: Claude Free, Pro, Max, Team and Enterprise
Frequently asked questions
How much does Claude Code cost per month?
The client is free to install, but Claude Code is not part of the Free plan. Access comes through Claude Pro at $20 a month ($17 billed annually), Max from $100 a month, a Team or Enterprise seat, or the Anthropic API billed per token with no minimum. Anthropic's cost documentation puts average enterprise spend at about $13 per developer per active day and $150 to $250 per month.
Can Claude Code run unattended in CI?
Yes, with the -p flag in non-interactive mode and the dontAsk permission mode, which denies every tool call that would otherwise prompt. The documentation is explicit that any session started with --dangerously-skip-permissions belongs inside a container, virtual machine or the sandbox runtime, running as a non-root user.
What does a PreToolUse hook actually enforce?
It runs before the tool call, in every permission mode including dontAsk and bypassPermissions, and can return permissionDecision deny to block it. A deny from a hook still blocks the call in bypassPermissions mode, but an allow from a hook cannot override a deny rule that already exists in settings.
Does the sandbox cover MCP servers and hooks?
No. The Bash sandbox applies to Bash, PowerShell and Monitor commands and their child processes on macOS, Linux and WSL2. Read, Edit, WebFetch, MCP servers, command hooks, plugin monitors and language servers all run outside that boundary, and native Windows runs unsandboxed.