Tools/AI agents

OpenAI Codex CLI: the agent that treats permissions as a config file

Codex CLI is OpenAI's open-source terminal coding agent. How its sandbox, approval policy and config.toml shape up, and what it costs to run unattended in CI.

Type
Coding agent
Pricing
Included with ChatGPT plans · API pay per token

··11 min read

  • Coding agent
  • Terminal
  • Sandbox
  • CI
  • Rust
A terminal session showing a prompt going to the model, a sandboxed shell command, and an approval prompt before the command runs.

Key takeaways

  • Codex CLI is Apache-2.0 licensed, written largely in Rust, and shipped as versioned npm packages, so it can be pinned in CI rather than curled at runtime.
  • The permission model is two independent settings, sandbox_mode and approval_policy, and changing who reviews an approval never widens the sandbox.
  • codex exec runs read-only by default, streams progress on stderr and prints only the final message on stdout, which makes it safe to wire into a pipeline.
  • The CLI requires a Git repository and refuses to run outside one unless you pass --skip-git-repo-check.
  • Signed-in ChatGPT plans and API-key billing are different accounting regimes: the subscription meter is message-shaped, the API is token-shaped, and they cannot be estimated against each other.

OpenAI Codex CLI is a terminal coding agent: it reads a repository, edits files, runs the project's own commands and iterates until the task is done. The interesting part is not the agent loop, which several competitors also do. It is that the safety model is expressed as two settings in a TOML file that a team can commit, review and pin, rather than as prompts or as a policy the vendor applies at the account level.

That makes it the most governable coding agent in the category, and it also makes it the one where configuration mistakes are most expensive, because a committed sandbox_mode = "danger-full-access" is a policy decision that reaches production without anybody reviewing a diff. Everything below is an assessment of how well that trade-off holds up.

What it is

The CLI is published as versioned npm packages, currently 0.160.1, and the repository is Apache-2.0 with a Rust core. It is one surface of a wider Codex product that also includes the ChatGPT desktop app, an IDE extension and a cloud runner, but the CLI is the part that runs on a developer's machine and the only part that is open source.

  • Apache-2.0 licensed, Rust core, npm package @openai/codex plus a standalone install script.
  • Configuration lives in ~/.codex/config.toml and can be scoped per project in .codex/config.toml.
  • Two separate sign-in paths: ChatGPT subscription or an API key, with different limits and different billing.
  • Local work needs a Git repository; codex exec enforces the same rule and offers --skip-git-repo-check.
  • MCP servers are configured in the same config file and shared with the IDE extension and the desktop app on the same host.

How it works

The turn structure is the familiar one: a prompt plus repository context go to a model, the model emits tool calls, the CLI runs them inside a sandbox, and the results come back as tool output for the next turn. What differs from a shell wrapper is that the sandbox and the approval check are separate gates, and both are configured rather than prompted for.

One Codex turn, with the sandbox and the approval gateThe prompt and the AGENTS.md instructions go to the model. The model emits a tool call. The sandbox decides whether the call touches only the workspace, and the approval policy decides whether it pauses for the user first. If the call runs, its output returns to the model and the loop continues.prompttask, AGENTS.mdmodeltool callsandboxworkspace-writehostfiles, git, commandsapprovalon-request1 send2 call3 allowed4 output5 next turn
The sandbox and the approval check are separate. Widening one does not widen the other.

The loop is not the product, though. What a team actually configures is the boundary around it, and the documentation is unusually explicit that the two are orthogonal. A run in Full access edits any file on the machine and runs commands with the network without asking, and the docs describe that as a significant increase in the risk of data loss and leaks rather than burying it.

Getting started

Install, sign in, run. The first launch offers Sign in with ChatGPT or an API key, and the choice matters later because it decides which limits apply and which billing regime you are in.

# Install on macOS or Linux; npm and Homebrew are also supported.
curl -fsSL https://chatgpt.com/codex/install.sh | sh

# Run inside a project directory, then sign in.
codex

# Pin the version in CI instead of tracking latest.
npm install -g @openai/codex@0.160.1

# The first prompt is a good place for /init, which writes AGENTS.md.
# /status, /model, /permissions and /review are the other useful commands.

AGENTS.md is the durable instruction file. The CLI reads a global ~/.codex/AGENTS.md, then walks from the project root down to the current directory, taking one file per level, with AGENTS.override.md winning over AGENTS.md. The combined size is capped by project_doc_max_bytes, which defaults to 32 KiB. For a team, that file is the most valuable thing to write, because it is the part of the agent's behaviour that goes through code review.

Sandbox and approvals

This is the part worth reading twice. The sandbox is the boundary; the approval policy is the pause. The default, Ask for approval, is sandbox_mode = "workspace-write" with approval_policy = "on-request" and a human reviewer. The docs are careful to note that switching the reviewer to auto_review does not widen the sandbox, which is the correct design and also the detail most tools get wrong.

# ~/.codex/config.toml
model = "gpt-6.1-sol"
model_reasoning_effort = "medium"

# Boundary: read-only | workspace-write | danger-full-access
sandbox_mode = "read-only"

# Pausing: on-request | never | granular table
approval_policy = "on-request"

# Reviewer: user | auto_review
approvals_reviewer = "user"

# Extra roots and network, only used when sandbox_mode = workspace-write
[sandbox_workspace_write]
writable_roots = ["~/code"]
network_access = false
SettingValuesWhat it does
sandbox_moderead-only, workspace-write, danger-full-accessWhich files and network the agent can reach. Defaults to read-only.
approval_policyon-request, never, granularWhen the agent pauses. Defaults to on-request.
approvals_revieweruser, auto_reviewWho answers a prompt. Does not change the sandbox.
/permissionsAsk for approval, Approve for me, Full accessInteractive presets over the same two settings.

The config file goes deeper than this. There is a network proxy with per-domain allow and deny rules, a shell environment policy that filters variables whose names look like keys or tokens, writable roots, and hooks that run before a tool call. That is a lot of surface, and it is worth resisting the urge to configure all of it: every knob is a decision somebody has to make on someone else's behalf.

Automation and codex exec

The non-interactive mode is the reason to pick this CLI over a chat-driven agent. codex exec runs in a read-only sandbox by default, streams progress to stderr, prints only the final message to stdout, and refuses to start outside a Git repository. That last rule is a guardrail against an agent rewriting a directory it cannot show a diff for.

# Progress on stderr, final message on stdout: safe to pipe.
codex exec "summarise the repository structure" | tee summary.md

# Machine-readable: one JSON object per event.
codex exec --json "triage the open bug reports" | jq

# Escalate the sandbox explicitly, never implicitly.
codex exec --sandbox workspace-write "add a regression test and run it"

# Structured final answer against a schema, written to a file.
codex exec "extract project metadata" \
  --output-schema ./schema.json -o ./metadata.json

Two details make it production-shaped. With --json the event stream carries usage, including cached input tokens, so a pipeline can attribute cost per run rather than per month. And an MCP server marked required = true fails startup instead of silently continuing without it, which is the difference between a broken CI run and a subtly wrong one.

MCP and project context

MCP servers are configured in the same config file, so they are shared across the CLI, the IDE extension and the desktop app on the same host. Both stdio and streamable HTTP transports are supported, with OAuth including dynamic client registration.

[mcp_servers.docs]
command = "npx"
args = ["-y", "@upstash/context7-mcp"]
env_vars = ["CONTEXT7_TOKEN"]
required = true          # fail startup instead of running without it
enabled_tools = ["search", "summarize"]
default_tools_approval_mode = "prompt"
tool_timeout_sec = 45

The cost of MCP is context, and the pricing documentation says so plainly: every server adds to the message and uses more of the allowance. A sensible team keeps two or three, disables the rest, and treats the server list as part of the per-turn budget rather than as a convenience list.

Cost and limits

Two accounting regimes share the same binary, and they are not comparable. A ChatGPT plan meters messages in five-hour windows, and the published figures are estimates rather than caps: roughly 15 to 160 local messages per five hours on the Plus plan for the mid-tier models, against 350 to 3,000 on the cheapest. An API key meters tokens at published rates, which is what makes it the right credential for automation.

The practical advice is short. Use the subscription for interactive work where a human is present to notice when the allowance is running out, and the API key for anything unattended, because only the second produces a predictable line item per run. The /status command shows remaining capacity in-session, and the docs are clear that prompt length alone is not a reliable predictor of what a task will consume.

Where it falls short

Three problems, in order of how much they matter. First, the configuration surface is enormous for a tool whose core loop is unremarkable, and there is no way to run it with a deliberately small config that you can audit in one sitting. Second, the product around the CLI churns fast: models are deprecated on fixed dates, several were retired inside a single year, and a committed model name in a config file or a CI script becomes a liability. Third, the code review command reports findings without editing the tree, which is the right behaviour and much less useful than the ones that fix what they find.

Codex CLIClaude CodeCline CLI
Non-interactive entry pointcodex execclaude -pcline --json
Machine-readable outputJSONL events, JSON Schema outputJSON or stream-json, --json-schemaNewline-delimited messages
Automation defaultread-only sandboxBare mode skips local configauto-approve true
ConfigurationOne TOML file, shared across clientsSettings files, MCP config, hooksConfig view and CLI flags

The comparison is close enough that the deciding factor is rarely the agent. It is what your permissions story already looks like. If the team already has a committed agent configuration under review, Codex fits. If not, the CLI's strength is a liability until someone writes that file.

Verdict

Codex CLI is the strongest choice for teams that want a coding agent whose behaviour is a reviewable artefact. The sandbox and approval split is well designed, the read-only default for automation is the right default, and being Apache-2.0 means the configuration can be vendored and pinned. It is a weaker choice for a single developer who wants to try an agent without first writing a policy.

  1. Adopt it when the agent's permissions should live in a file that goes through code review, and pin the npm version in CI.
  2. Use an API key for anything unattended. Subscription credit accounting does not survive contact with a scheduled job.
  3. Write AGENTS.md before tuning anything else. It changes behaviour more than any flag in config.toml.
  4. Keep sandbox_mode at workspace-write for automation, and reserve danger-full-access for a container you control.
  5. Skip it if you need a stable model identifier over a year. The deprecation cadence is faster than most teams can absorb.

Sources

  1. Codex CLI documentation
  2. Codex: configuration
  3. Codex: sample configuration
  4. Codex: permissions
  5. Codex: non-interactive mode
  6. Codex: authentication
  7. Codex: Model Context Protocol
  8. Codex: pricing
  9. Codex: open-source components
  10. Codex changelog

Frequently asked questions

Is Codex CLI open source?

Yes. The CLI, the SDK and the app server live in the openai/codex repository under the Apache-2.0 licence. The IDE extension and Codex Cloud are not open source, and the security CLI ships separately as openai/codex-security.

How do I run Codex CLI in CI without it touching anything?

Use codex exec, which runs in a read-only sandbox unless you pass --sandbox workspace-write. Progress goes to stderr and the final agent message goes to stdout, so a pipeline can capture the answer without parsing progress lines. An API key is the right credential in CI, because the API bills per token.

What is the difference between sandbox_mode and approval_policy?

sandbox_mode sets the boundary: read-only, workspace-write or danger-full-access. approval_policy sets when the agent pauses: on-request, never, or a granular table. They are separate, and switching the reviewer from user to auto_review keeps the same sandbox boundary.

Does Codex CLI work without a Git repository?

Not by default. The CLI requires commands to run inside a Git repository to prevent destructive changes, and codex exec refuses to start otherwise. --skip-git-repo-check overrides it, which is only reasonable in a container you already control.

Sounds like what you need?

Tell me about your project or role – I’d love to hear from you.