Tools/AI agents
OpenAI Codex CLI: the agent that treats permissions as a config file
Codex CLI is OpenAI's open-source terminal coding agent. How its sandbox, approval policy and config.toml shape up, and what it costs to run unattended in CI.
- Type
- Coding agent
- Pricing
- Included with ChatGPT plans · API pay per token
Balázs Csorba··11 min read
- Coding agent
- Terminal
- Sandbox
- CI
- Rust

Key takeaways
- Codex CLI is Apache-2.0 licensed, written largely in Rust, and shipped as versioned npm packages, so it can be pinned in CI rather than curled at runtime.
- The permission model is two independent settings, sandbox_mode and approval_policy, and changing who reviews an approval never widens the sandbox.
- codex exec runs read-only by default, streams progress on stderr and prints only the final message on stdout, which makes it safe to wire into a pipeline.
- The CLI requires a Git repository and refuses to run outside one unless you pass --skip-git-repo-check.
- Signed-in ChatGPT plans and API-key billing are different accounting regimes: the subscription meter is message-shaped, the API is token-shaped, and they cannot be estimated against each other.
OpenAI Codex CLI is a terminal coding agent: it reads a repository, edits files, runs the project's own commands and iterates until the task is done. The interesting part is not the agent loop, which several competitors also do. It is that the safety model is expressed as two settings in a TOML file that a team can commit, review and pin, rather than as prompts or as a policy the vendor applies at the account level.
That makes it the most governable coding agent in the category, and it also makes it the one where configuration mistakes are most expensive, because a committed sandbox_mode = "danger-full-access" is a policy decision that reaches production without anybody reviewing a diff. Everything below is an assessment of how well that trade-off holds up.
What it is
The CLI is published as versioned npm packages, currently 0.160.1, and the repository is Apache-2.0 with a Rust core. It is one surface of a wider Codex product that also includes the ChatGPT desktop app, an IDE extension and a cloud runner, but the CLI is the part that runs on a developer's machine and the only part that is open source.
- Apache-2.0 licensed, Rust core, npm package
@openai/codexplus a standalone install script. - Configuration lives in
~/.codex/config.tomland can be scoped per project in.codex/config.toml. - Two separate sign-in paths: ChatGPT subscription or an API key, with different limits and different billing.
- Local work needs a Git repository;
codex execenforces the same rule and offers--skip-git-repo-check. - MCP servers are configured in the same config file and shared with the IDE extension and the desktop app on the same host.
How it works
The turn structure is the familiar one: a prompt plus repository context go to a model, the model emits tool calls, the CLI runs them inside a sandbox, and the results come back as tool output for the next turn. What differs from a shell wrapper is that the sandbox and the approval check are separate gates, and both are configured rather than prompted for.
The loop is not the product, though. What a team actually configures is the boundary around it, and the documentation is unusually explicit that the two are orthogonal. A run in Full access edits any file on the machine and runs commands with the network without asking, and the docs describe that as a significant increase in the risk of data loss and leaks rather than burying it.
Getting started
Install, sign in, run. The first launch offers Sign in with ChatGPT or an API key, and the choice matters later because it decides which limits apply and which billing regime you are in.
# Install on macOS or Linux; npm and Homebrew are also supported.
curl -fsSL https://chatgpt.com/codex/install.sh | sh
# Run inside a project directory, then sign in.
codex
# Pin the version in CI instead of tracking latest.
npm install -g @openai/codex@0.160.1
# The first prompt is a good place for /init, which writes AGENTS.md.
# /status, /model, /permissions and /review are the other useful commands.AGENTS.md is the durable instruction file. The CLI reads a global ~/.codex/AGENTS.md, then walks from the project root down to the current directory, taking one file per level, with AGENTS.override.md winning over AGENTS.md. The combined size is capped by project_doc_max_bytes, which defaults to 32 KiB. For a team, that file is the most valuable thing to write, because it is the part of the agent's behaviour that goes through code review.
Sandbox and approvals
This is the part worth reading twice. The sandbox is the boundary; the approval policy is the pause. The default, Ask for approval, is sandbox_mode = "workspace-write" with approval_policy = "on-request" and a human reviewer. The docs are careful to note that switching the reviewer to auto_review does not widen the sandbox, which is the correct design and also the detail most tools get wrong.
# ~/.codex/config.toml
model = "gpt-6.1-sol"
model_reasoning_effort = "medium"
# Boundary: read-only | workspace-write | danger-full-access
sandbox_mode = "read-only"
# Pausing: on-request | never | granular table
approval_policy = "on-request"
# Reviewer: user | auto_review
approvals_reviewer = "user"
# Extra roots and network, only used when sandbox_mode = workspace-write
[sandbox_workspace_write]
writable_roots = ["~/code"]
network_access = false| Setting | Values | What it does |
|---|---|---|
sandbox_mode | read-only, workspace-write, danger-full-access | Which files and network the agent can reach. Defaults to read-only. |
approval_policy | on-request, never, granular | When the agent pauses. Defaults to on-request. |
approvals_reviewer | user, auto_review | Who answers a prompt. Does not change the sandbox. |
/permissions | Ask for approval, Approve for me, Full access | Interactive presets over the same two settings. |
The config file goes deeper than this. There is a network proxy with per-domain allow and deny rules, a shell environment policy that filters variables whose names look like keys or tokens, writable roots, and hooks that run before a tool call. That is a lot of surface, and it is worth resisting the urge to configure all of it: every knob is a decision somebody has to make on someone else's behalf.
Automation and codex exec
The non-interactive mode is the reason to pick this CLI over a chat-driven agent. codex exec runs in a read-only sandbox by default, streams progress to stderr, prints only the final message to stdout, and refuses to start outside a Git repository. That last rule is a guardrail against an agent rewriting a directory it cannot show a diff for.
# Progress on stderr, final message on stdout: safe to pipe.
codex exec "summarise the repository structure" | tee summary.md
# Machine-readable: one JSON object per event.
codex exec --json "triage the open bug reports" | jq
# Escalate the sandbox explicitly, never implicitly.
codex exec --sandbox workspace-write "add a regression test and run it"
# Structured final answer against a schema, written to a file.
codex exec "extract project metadata" \
--output-schema ./schema.json -o ./metadata.jsonTwo details make it production-shaped. With --json the event stream carries usage, including cached input tokens, so a pipeline can attribute cost per run rather than per month. And an MCP server marked required = true fails startup instead of silently continuing without it, which is the difference between a broken CI run and a subtly wrong one.
MCP and project context
MCP servers are configured in the same config file, so they are shared across the CLI, the IDE extension and the desktop app on the same host. Both stdio and streamable HTTP transports are supported, with OAuth including dynamic client registration.
[mcp_servers.docs]
command = "npx"
args = ["-y", "@upstash/context7-mcp"]
env_vars = ["CONTEXT7_TOKEN"]
required = true # fail startup instead of running without it
enabled_tools = ["search", "summarize"]
default_tools_approval_mode = "prompt"
tool_timeout_sec = 45The cost of MCP is context, and the pricing documentation says so plainly: every server adds to the message and uses more of the allowance. A sensible team keeps two or three, disables the rest, and treats the server list as part of the per-turn budget rather than as a convenience list.
Cost and limits
Two accounting regimes share the same binary, and they are not comparable. A ChatGPT plan meters messages in five-hour windows, and the published figures are estimates rather than caps: roughly 15 to 160 local messages per five hours on the Plus plan for the mid-tier models, against 350 to 3,000 on the cheapest. An API key meters tokens at published rates, which is what makes it the right credential for automation.
The practical advice is short. Use the subscription for interactive work where a human is present to notice when the allowance is running out, and the API key for anything unattended, because only the second produces a predictable line item per run. The /status command shows remaining capacity in-session, and the docs are clear that prompt length alone is not a reliable predictor of what a task will consume.
Where it falls short
Three problems, in order of how much they matter. First, the configuration surface is enormous for a tool whose core loop is unremarkable, and there is no way to run it with a deliberately small config that you can audit in one sitting. Second, the product around the CLI churns fast: models are deprecated on fixed dates, several were retired inside a single year, and a committed model name in a config file or a CI script becomes a liability. Third, the code review command reports findings without editing the tree, which is the right behaviour and much less useful than the ones that fix what they find.
| Codex CLI | Claude Code | Cline CLI | |
|---|---|---|---|
| Non-interactive entry point | codex exec | claude -p | cline --json |
| Machine-readable output | JSONL events, JSON Schema output | JSON or stream-json, --json-schema | Newline-delimited messages |
| Automation default | read-only sandbox | Bare mode skips local config | auto-approve true |
| Configuration | One TOML file, shared across clients | Settings files, MCP config, hooks | Config view and CLI flags |
The comparison is close enough that the deciding factor is rarely the agent. It is what your permissions story already looks like. If the team already has a committed agent configuration under review, Codex fits. If not, the CLI's strength is a liability until someone writes that file.
Verdict
Codex CLI is the strongest choice for teams that want a coding agent whose behaviour is a reviewable artefact. The sandbox and approval split is well designed, the read-only default for automation is the right default, and being Apache-2.0 means the configuration can be vendored and pinned. It is a weaker choice for a single developer who wants to try an agent without first writing a policy.
- Adopt it when the agent's permissions should live in a file that goes through code review, and pin the npm version in CI.
- Use an API key for anything unattended. Subscription credit accounting does not survive contact with a scheduled job.
- Write AGENTS.md before tuning anything else. It changes behaviour more than any flag in config.toml.
- Keep sandbox_mode at workspace-write for automation, and reserve danger-full-access for a container you control.
- Skip it if you need a stable model identifier over a year. The deprecation cadence is faster than most teams can absorb.
Sources
Frequently asked questions
Is Codex CLI open source?
Yes. The CLI, the SDK and the app server live in the openai/codex repository under the Apache-2.0 licence. The IDE extension and Codex Cloud are not open source, and the security CLI ships separately as openai/codex-security.
How do I run Codex CLI in CI without it touching anything?
Use codex exec, which runs in a read-only sandbox unless you pass --sandbox workspace-write. Progress goes to stderr and the final agent message goes to stdout, so a pipeline can capture the answer without parsing progress lines. An API key is the right credential in CI, because the API bills per token.
What is the difference between sandbox_mode and approval_policy?
sandbox_mode sets the boundary: read-only, workspace-write or danger-full-access. approval_policy sets when the agent pauses: on-request, never, or a granular table. They are separate, and switching the reviewer from user to auto_review keeps the same sandbox boundary.
Does Codex CLI work without a Git repository?
Not by default. The CLI requires commands to run inside a Git repository to prevent destructive changes, and codex exec refuses to start otherwise. --skip-git-repo-check overrides it, which is only reasonable in a container you already control.