Tools/AI agents

Claude Code: the terminal coding agent, reviewed

An engineering review of Claude Code: the extension surface, the real cost per developer, and the exact boundary of the Bash sandbox.

Type
Coding agent
Pricing
Free · pay per API token

··10 min read

  • Terminal agent
  • Hooks
  • Subagents
  • MCP
  • Sandbox
Cover art for the Claude Code review: a terminal session feeding a permission gate, a context window and a sandbox boundary

Key takeaways

  • Version 2.1.292 shipped on 6 October 2026 and the changelog lists releases almost daily, so pinning a version for CI is standing work rather than a one-off.
  • Six mechanisms extend the loop — CLAUDE.md, skills, subagents, MCP, hooks and plugins — and using the wrong one is the most common configuration mistake.
  • Permission rules are enforced by the client, not the model, which makes a PreToolUse hook the only guard that still holds in bypassPermissions.
  • The Bash sandbox covers Bash, PowerShell and Monitor commands only; file tools, MCP servers, hooks and language servers all run outside it.
  • Anthropic's own cost documentation puts average enterprise spend at about $13 per developer per active day, with 90% of users below $30.

Claude Code is Anthropic's terminal-first coding agent: a CLI that reads a repository, edits files, runs commands and iterates on the result, with the Claude models behind it. On the evidence of the documentation it is the most complete agent harness on the market, and the reason is not the model but the six extension mechanisms wrapped around the loop. The cost of that completeness is a configuration surface large enough to get wrong, and a release cadence — 2.1.292 landed on 6 October 2026 — that turns version pinning into standing maintenance. Worth adopting, provided the security section below is read before the first unattended run.

It competes with Cursor's editor-native agent, GitHub Copilot's agent mode and the Codex CLI, and it is the only one of the four where the agent is the application rather than a feature of an editor. That matters on a large repository, where the agent gets the whole working tree, the shell and the git history rather than whatever happens to be open in a tab. The trade is the interface: a terminal is a poor place to read a diff, which is why the VS Code and JetBrains extensions exist at all.

What it actually is

The agentic loop is the product. A task moves through gather context, take action, verify results, over and over, with the model choosing the next tool call and the harness supplying the tools: file operations, search, shell execution, web search, and code intelligence through language-server plugins. Everything below the loop — settings files, skills, subagents, hooks — exists to make that loop cheaper, safer or broader.

  • Version 2.1.292, published 6 October 2026; the changelog lists releases up to that point almost daily.
  • Runs in the terminal, as VS Code and JetBrains extensions, on the desktop, in the browser, and on Amazon Bedrock and Google Cloud.
  • Built-in tools cover file operations, search, shell, web search and code intelligence; the terminal CLI, VS Code and JetBrains also accept third-party providers.
  • Permission modes are default (labelled manual), acceptEdits, plan, auto, dontAsk and bypassPermissions.
  • Auto mode is the built-in starting mode for interactive terminal and VS Code sessions on Pro, Max and Team: a separate classifier model reviews each action instead of a human.
  • The Bash sandbox uses operating-system primitives — Seatbelt on macOS, bubblewrap on Linux — and applies to Bash, PowerShell and Monitor commands only.
  • Bundled skills ship in the box, including /doctor, /code-review, /batch, /debug and /loop.

How it works: one context window

Everything the model knows about a session lives in a single context window: CLAUDE.md and any AGENTS.md, auto memory, MCP tool names, skill descriptions, every file read, every tool result and the transcript itself. Path-scoped rules load when their trigger file is read. When the window fills, the session compacts — the conversation is replaced with a structured summary, the root CLAUDE.md and unscoped rules are re-injected from disk, up to five recently modified files are re-read, and invoked skill bodies come back capped at 5,000 tokens each and 25,000 in total.

Where a session's tokens go, and what compaction keepsFour columns. Startup: CLAUDE.md, unscoped rules, auto memory and skill names, re-injected from disk after compaction. Reads: file contents, tool output and web results, kept only as part of the summary. History: the transcript and its thinking, replaced by a structured summary. Output: the final text, the diffs and the summaries returned by subagents, which are not summarised away. A band below: after compaction the project CLAUDE.md, unscoped rules, auto memory, the git status snapshot and the plan are re-injected from disk, while up to five recently modified files, path-scoped rules and nested CLAUDE.md files are rebuilt on demand. Invoked skill bodies return capped at 5,000 tokens each and 25,000 in total.Where a session’s tokens goand what compaction keepsStartuploaded before you typeCLAUDE.mdunscoped rulesauto memoryskill namesReadsgrows with the taskfile contentstool outputweb resultssearch hitsHistorygrows every turntranscriptthinking blockstool call resultsre-sent each turnOutputwhat comes backfinal textdiffssubagent summariesnot summarisedWhen the window fills: compactionRe-injected from disk: project CLAUDE.md, unscopedrules, auto memory, a fresh git status, the planRebuilt on demand: up to five recently modifiedfiles, path-scoped rules, nested CLAUDE.mdskill bodies return capped at 5,000 tokens each and 25,000 in total
Context is the scarce resource: everything the model knows about a session is re-sent on every turn, and compaction is the mechanism that throws part of it away.

The consequence for a bill is that re-sent history dominates, not the answer. A real session screen reports hundreds of thousands of cached input tokens against a few thousand output tokens, and cache reads are billed at a tenth of the input rate on current models. Anyone trying to reduce spend is really deciding what enters the window, which is why the documentation pushes MCP Tool Search, subagents and skills over simply asking the model to be brief.

The extension surface

Six mechanisms extend the loop, and the documentation is unusually clear that they are not interchangeable. CLAUDE.md is always-on context. A skill is knowledge or a workflow loaded on demand, following the Agent Skills open standard. A subagent is an isolated context that returns a summary instead of its transcript. MCP connects external services. A hook is a shell command, HTTP request, MCP tool call, single-turn prompt or agent that fires on a lifecycle event. Plugins bundle the lot for distribution across a team.

MechanismWhat it isContext costDeterministic?
CLAUDE.mdAlways-on project instructionsInjected at session startNo, the model decides to follow it
SkillMarkdown knowledge or workflowDescription at start, body on useNo
SubagentIsolated loop returning a summaryOnly the summary returnsNo
MCP serverExternal tools and dataNames at start, schemas on useNo
HookScript or model call on an eventZero unless it returns outputYes, the event always fires
PluginBundle of skills, hooks, agents, MCPWhatever the bundle containsDepends on contents

One sentence from the permissions documentation decides the whole design: permission rules are enforced by Claude Code, not by the model. An instruction like never edit .env in CLAUDE.md is a request; a PreToolUse hook that denies the edit is enforcement. Any team treating the first as a control has misread the threat model.

Getting started properly

Installation is a shell script and a login; there is no project to set up. The part of a first session worth getting right is the settings file, because that is where a team converts prompt instructions into enforced rules. This project-level configuration allowlists the commands it trusts, blocks a git push, formats every edit and switches on the Bash sandbox:

{
  "permissions": {
    "allow": [
      "Bash(npm run *)",
      "Bash(git commit *)",
      "Read",
      "Edit(src/**)"
    ],
    "deny": [
      "Bash(git push *)",
      "Read(.env)"
    ]
  },
  "sandbox": {
    "enabled": true,
    "allowWrite": ["src"],
    "allowedDomains": ["registry.npmjs.org"]
  },
  "hooks": {
    "PostToolUse": [
      {
        "matcher": "Edit|Write",
        "hooks": [
          { "type": "command", "command": "jq -r '.tool_input.file_path' | xargs npx prettier --write" }
        ]
      }
    ]
  }
}

Two details in that file are easy to get wrong. Bash(git commit *) matches git commit -m 'x' but not git -C . commit -m 'x', because the wildcard stands in for whatever text sits in its place. And a hook that exits 0 with no output has not approved anything — it has declined to decide, and the call continues through the normal permission flow. What does hold everywhere is a deny returned by a PreToolUse hook: it still blocks the tool in bypassPermissions mode, which is the one guarantee that survives a developer who has switched their own prompts off.

What it costs

The client is free to install and Claude Code is not part of the Free plan. Model access comes from a Claude subscription, a Team or Enterprise seat, or the Anthropic API billed per token. What makes the bill hard to predict is that subscriptions meter against rolling usage windows rather than publishing a token allowance.

RoutePriceClaude CodeShape
Free$0Not includedChat only, no terminal agent
Pro$20 a month, $17 billed annuallyYesRolling session and weekly limits
MaxFrom $100 a monthYes5x or 20x Pro usage
Team standard seat$20 a seat annually, $25 monthlyYesMix and match with premium seats
Team premium seat$100 a seat annually, $125 monthlyYesFive times the standard seat usage
Enterprise$20 a seat annually plus usage at API ratesYesSSO, SCIM, audit logs, spend limits
Anthropic APIPer token, no minimumYesHard ceiling, billed directly

On the per-token route the current list prices are $4 per million input tokens and $20 per million output tokens for Opus 5.5, $2 and $10 for Sonnet 5.5, and $1 and $5 for Haiku 4.5, with cache reads at $0.20, $0.20 and $0.10. The cheapest lever is therefore the model: the same task on Sonnet 5.5 instead of Opus 5.5 halves the bill at list price, and Haiku 4.5 is a quarter of it. The documentation pushes the same point from the other end — moving long instructions out of CLAUDE.md and into skills, keeping MCP servers few, and pushing verbose work into subagents so their transcripts never enter the main window.

Security and isolation

The security model is the strongest part of the design and the easiest to misuse. Rules are matched by Claude Code rather than by the model, precedence puts a deny at any scope above an allow at any other, and managed settings outrank user settings, project settings and command-line flags. Deny rules hold in bypassPermissions, and a removal such as rm -rf / is refused even when an allow rule or a hook permits it. That is the behaviour you want from a circuit breaker.

The sandbox is the second layer, and its scope is narrower than the name suggests:

  • Bash, PowerShell and Monitor commands and their child processes, on macOS, Linux and WSL2 — native Windows runs unsandboxed.
  • File tools such as Read, Edit, Write and WebFetch, which follow permission rules instead; a sandbox denyRead entry does not stop Read.
  • MCP servers, command hooks, plugin monitors, language servers and helper commands, all of which run with the session's full access.
  • Commands that fail inside the sandbox can be retried unsandboxed through dangerouslyDisableSandbox; setting allowUnsandboxedCommands to false removes the escape hatch.
  • Trust is per directory and lasts one session: a project-level subagent's frontmatter hooks do not run until the workspace trust dialog is accepted, and a -p run does not count as accepting it.

Where it weakens

The honest weaknesses come before the comparison. Claude Code is fast-moving and configuration-heavy, the model behind it is not yours to tune, and the agent has shell access by design.

  • A release a day means behaviour can change under a CI job. --bare exists to strip hooks, skills, commands, subagents, plugins, MCP servers, auto memory and CLAUDE.md for reproducible scripted runs, and it is the right default there.
  • Context is the scarce resource. Compaction drops the middle of the conversation, only five files are re-read afterwards, and skill bodies are truncated from the start — so the important instruction at the bottom of a long SKILL.md is the one that disappears.
  • Descriptions of model-invocable skills load on every request, so vague or overlapping descriptions make the model load the wrong skill or miss the useful one.
  • Subscriptions meter in time windows, not dollars. One heavy morning can exhaust the session limit with a week of allowance still unused, and only API billing offers a hard ceiling.
  • It is closed. Anthropic's models are the product; only the terminal CLI, VS Code and JetBrains accept a third-party provider, and agent features are not portable to another harness.
OptionWhat it isEntry priceMain trade-off
Claude CodeTerminal agent, the application itself$20 a month on ProHighest ceiling for repository-scale work, largest configuration surface
CursorEditor fork with an agent mode$20 a month on ProBetter diff and inline ergonomics, but the agent lives inside the editor
GitHub CopilotIDE extension, completions plus agent$10 a month on ProCheapest entry point, least autonomy of the four
Codex CLITerminal agent on OpenAI modelsBundled with ChatGPT plansA different model family, without the Claude-specific harness features

The short version: for long-running, repository-scale work the terminal harness wins on capability; for line-by-line work an editor-native agent on a cheaper seat is the better buy. Running both is a defensible $30 a month, and most teams who end up there describe it as in-editor work for the small changes and the terminal for everything that spans a repository.

Verdict

Claude Code is worth adopting on one condition: the team writes the guardrails down and puts them in version control. Without a committed settings file, a PreToolUse hook on the protected paths and an explicit decision about unattended runs, the tool is faster than a reviewer and less careful than one.

  1. Adopt it when the work is repository-scale: migrations, refactors, multi-file changes that end in a verification step.
  2. Adopt it when the budget can carry one seat for a heavy user and one for a light user — the windows, not the seats, are the binding constraint.
  3. Adopt the permission and hook model seriously. It is the only difference between an agent and a supervised agent.
  4. Do not make it the only tool. Keep an editor-native completion product if most of the day is writing lines rather than changing systems.
  5. Do not run it unattended on a machine you care about without a container, and never with --dangerously-skip-permissions outside one.
Permission rules are enforced by Claude Code, not by the model. Instructions in your prompt or CLAUDE.md shape what Claude tries to do, but they do not change what Claude Code allows.

Sources

  1. Claude Code documentation: overview
  2. Claude Code documentation: how Claude Code works
  3. Claude Code documentation: extend Claude Code
  4. Claude Code documentation: hooks reference
  5. Claude Code documentation: configure permissions
  6. Claude Code documentation: sandboxing
  7. Claude Code documentation: manage costs effectively
  8. Claude Code documentation: explore the context window
  9. Claude Code changelog
  10. Claude Platform pricing: model list prices and prompt caching
  11. Anthropic pricing: Claude Free, Pro, Max, Team and Enterprise

Frequently asked questions

How much does Claude Code cost per month?

The client is free to install, but Claude Code is not part of the Free plan. Access comes through Claude Pro at $20 a month ($17 billed annually), Max from $100 a month, a Team or Enterprise seat, or the Anthropic API billed per token with no minimum. Anthropic's cost documentation puts average enterprise spend at about $13 per developer per active day and $150 to $250 per month.

Can Claude Code run unattended in CI?

Yes, with the -p flag in non-interactive mode and the dontAsk permission mode, which denies every tool call that would otherwise prompt. The documentation is explicit that any session started with --dangerously-skip-permissions belongs inside a container, virtual machine or the sandbox runtime, running as a non-root user.

What does a PreToolUse hook actually enforce?

It runs before the tool call, in every permission mode including dontAsk and bypassPermissions, and can return permissionDecision deny to block it. A deny from a hook still blocks the call in bypassPermissions mode, but an allow from a hook cannot override a deny rule that already exists in settings.

Does the sandbox cover MCP servers and hooks?

No. The Bash sandbox applies to Bash, PowerShell and Monitor commands and their child processes on macOS, Linux and WSL2. Read, Edit, WebFetch, MCP servers, command hooks, plugin monitors and language servers all run outside that boundary, and native Windows runs unsandboxed.

Sounds like what you need?

Tell me about your project or role – I’d love to hear from you.