Tools/AI agents

Goose: the open-source agent you run on your own machine

Goose is an Apache-2.0 coding agent written in Rust, now owned by the Linux Foundation. How its permission modes, recipes and MCP extensions hold up against commercial tools.

Type
Coding agent
Pricing
Free · self-host

··10 min read

  • Coding agent
  • MCP
  • Local models
  • Open source
The goose agent loop: prompt, goose core, model, tool call, MCP extension, running back to the prompt, with the permission mode, turn limit and compaction controls underneath.

Key takeaways

  • Goose is an Apache-2.0 coding agent in Rust, released by Block in January 2025 and donated to the Agentic AI Foundation at the Linux Foundation in December 2025, so governance no longer sits with its original sponsor.
  • Capability is a configuration decision: the tools a session can reach are the MCP extensions enabled for it, restricted per extension with an available_tools allowlist.
  • Four permission modes exist and the most permissive one, fully autonomous, is the default, so a fresh install edits and deletes files without asking.
  • Recipes are versioned YAML with parameters, retries and a machine-checkable success test, which makes a repeatable agent workflow something you can review and schedule rather than a prompt you retype.
  • The operational tax is real: most external extensions need npx or uvx on the machine, dev servers hang sessions until the tool times out, and the docs tell you to end and restart a session when the agent stops responding.

Goose is an open-source coding agent that runs on the developer's own machine rather than inside a vendor's cloud. It ships as a desktop app for macOS, Linux and Windows, a full CLI, and an API to embed, and it is written in Rust. Block released it publicly in January 2025 and donated it in December 2025 to the Agentic AI Foundation at the Linux Foundation, alongside Anthropic's Model Context Protocol and OpenAI's AGENTS.md, so the project is now governed by a foundation rather than by the company that started it.

What is not commodity here is the extension model. Goose reaches the outside world through Model Context Protocol servers, and the tools a session can use are the extensions enabled for it. That makes an agent's capability surface a configuration decision rather than a property of the binary, which is a materially different shape from an agent with a fixed built-in toolset.

The position after reading the documentation: this is a serious, well-documented agent that loses to the commercial leaders on polish and on individual task quality, and wins on the fact that every part of it can be inspected, forked and redistributed, including the format its workflows are written in. It is most convincing as a platform for standardising agent behaviour across a fleet, and least convincing as a single developer's daily driver when a commercial tool with better autocomplete is already licensed.

What goose is

Two surfaces on one core. The desktop app is the approachable half and the CLI is the scriptable half; both drive the same session logic, the same config file and the same extension list. The repository is Apache-2.0 and shows roughly 55,000 stars and 6,400 forks.

  • Apache-2.0 licensed and written in Rust, published as signed release artefacts and as a one-line install script for the CLI.
  • 15+ LLM providers including Anthropic, OpenAI, Google, Azure, Bedrock, OpenRouter and Ollama, with API keys or an existing subscription over ACP.
  • 70+ extensions over MCP, in four types: stdio, builtin, platform and streamable HTTP.
  • Four permission modes, from fully autonomous to chat only, chosen per session with granular per-tool permissions on top.
  • Recipes: versioned YAML files packaging instructions, extensions, parameters, settings, retries and a JSON response schema into one reviewable unit.
  • Auto-compaction at 80% of the context window, tool-output summarisation, and a hard turn ceiling that defaults to 1000.
  • A scheduler, an ACP server for editors such as Zed, a built-in review command and documented support for shipping your own branded build.

How it works

The loop is the standard one: send context, receive a message, and if it contains a tool call, execute the tool and feed the result into the next turn. What goose adds around that loop is bookkeeping. Token accounting, compaction, a turn limit, a cost estimate and a permission gate all sit between the model and the filesystem, and every one of them is a setting rather than a heuristic buried in the binary.

The goose agent loopA prompt enters the goose core, which calls a model, receives a tool call, runs it through an MCP extension, and loops back to the prompt. Three controls apply on every turn: the permission mode, the maximum turn limit and automatic context compaction.The goose agent looprust core, mcp toolsPromptsession, hintsgoose coresession loopModel15+ providersTool callshell, fileSESSION LOOPCONTROLSPermission modeauto, approve, chatMax turnsdefault 1000Auto-compactionat 80% of context
The agent loop is the commodity part. The three controls underneath it are the difference: each is an explicit setting that decides how far the agent may go unattended.

Sessions are one continuous conversation. A token indicator shows what is used against what is available, and once the compaction threshold is reached the older part is summarised while the previous messages stay visible in the interface. It is a sensible design, and it means a long session loses fidelity well before it fails.

# install the CLI (macOS, Linux), then pick a provider
curl -fsSL https://github.com/aaif-goose/goose/releases/download/stable/download_cli.sh | bash
goose configure

# run a session in a repository, with a turn cap
cd ~/code/my-service
goose session --max-turns 25 --with-builtin developer

# unattended: run a recipe and print only the model's reply
goose run --recipe release-check.yaml --params env=staging --quiet

# what provider, model and extensions are actually in play
goose info --verbose

Provider settings live in a YAML config file, with secrets in the system keyring by default. Setting GOOSE_DISABLE_KEYRING forces file storage instead, which is what the docs prescribe for containers and headless Linux where the keyring is often unavailable. Since version 1.10.0 sessions live in a SQLite database rather than a directory of JSONL files, so backups and migrations are a database problem.

Context, turns and cost

Three settings decide how long a session runs and what it costs, and all three are plain configuration keys rather than hidden behaviour. The docs even suggest values for each class of work, which is unusual and useful.

  • Auto-compaction threshold, 0.8 by default, the fraction of the context window at which old turns are summarised; set it to 0.0 to disable.
  • Maximum turns without human input, 1000 by default. On reaching it the agent stops and asks. The docs suggest 5 to 10 for exploratory work and 100 or more for a migration.
  • Prompt cache lifetime for Anthropic, 5 minutes by default or 1 hour to keep the cached prefix alive across idle gaps at a higher cache-write rate.
  • A per-session cost estimate priced from the OpenRouter catalogue and cached locally. The docs are explicit that it is a public-price estimate, not an invoice.

Context limits resolve from an explicit override, then declarative provider config, then runtime discovery, then model metadata, and finally a 128,000-token default. That last fallback matters when goose points at a gateway with a custom model name: the indicator shows the wrong number until the limit is set explicitly, which is a small trap for anyone running it against a proxy.

Extensions and recipes

An extension is an MCP server with a name, a command, a timeout and an optional list of the tools it may expose. That list is the setting that matters most in practice: every tool a session can see is a choice the model has to make correctly on each turn, and narrowing the list is cheaper than a better prompt.

version: "1.0.0"
title: "Release risk check"
description: "Diff the branch against a base and report shipping risk"
parameters:
  - key: base
    input_type: string
    requirement: required
    description: "Branch to compare against"
prompt: |
  Compare this branch against {{ base }} and summarise
  behavioural risk in shipping order.

extensions:
  - type: builtin
    name: developer
    timeout: 300
  - type: stdio
    name: github
    cmd: npx
    args: ["-y", "@modelcontextprotocol/server-github"]
    env_keys: [GITHUB_PERSONAL_ACCESS_TOKEN]
    available_tools: [get_file_contents, search_code]

settings:
  goose_provider: anthropic
  temperature: 0.2
  max_turns: 40

retry:
  max_retries: 2
  timeout_seconds: 30
  checks:
    - type: shell
      command: "test -f RISK.md"

That is the shape worth copying. A recipe is reviewable YAML with a validation command attached, so a scheduled job can be retried until it produces something checkable instead of retried blindly, and validation rejects a template variable with no matching parameter definition. Sub-recipes compose the same way, with the parent mapping its values onto the children, though the docs still mark that feature experimental and note that sub-recipes run in isolated sessions with no shared memory.

Permissions and autonomy

Permission is the whole safety story, and the default is the most permissive setting available. There are four modes, and each session picks its own.

  • Autonomous, which modifies files, uses extensions and deletes files with no approval. This is the default.
  • Manual approval, which asks before any tool and honours the granular per-tool permission list.
  • Smart approval, which auto-approves what it reads as low risk and flags the rest. The classification is performed by the model provider, so it is a suggestion rather than a control.
  • Chat only, which forbids extensions and file changes, for analysis, writing and reasoning.

Where it shingles

The gaps are practical rather than architectural. Most external extensions launch through npx or uvx, so the agent quietly depends on Node.js or a Python toolchain being installed. It will start development servers nobody asked for, and those never exit, so the session hangs until the tool times out. Extensions download their runtimes through a bundled copy of Hermit, which fails on an air-gapped network unless the shims are renamed. Windows installs expect Node.js at a fixed path and the documented remedy is a symbolic link. And the known-issues page advises ending the session and starting a new one when the agent stops responding, which is honest but not a fix.

gooseClaude CodeOpenHands
Where it runsDesktop, CLI or embedded APITerminal onlyLocal or Docker, with a web UI
LicenceApache-2.0ProprietaryOpen source
ToolsMCP extensions you enableBuilt in, plus MCPBuilt in, plus MCP
Model choice15+ providers or your own gatewayAnthropic onlyBring your own key
Weakest pointA runtime on every machineOne vendor, one price listHeavier runtime than a CLI

The trade against the commercial tools is straightforward. Claude Code is a better daily driver: it starts faster, the model quality is the reason people stay, and its permission model comes from the provider instead of being inferred by it. OpenHands wins exactly where goose is weakest, running agent work in a sandbox where the result can be inspected, at the cost of a heavier runtime. Goose wins the two things that matter to a platform team: a licence that permits forking, and a declarative format with an allowlist per recipe that can be enforced across a fleet.

The second structural gap is security depth. There is prompt-injection detection with a configurable threshold, and an adversary mode that runs a separate reviewer agent over tool calls, which is an interesting idea and not a control anyone should rely on alone. Read-only tools execute without approval by default, and extensions run as the developer with the developer's credentials.

Verdict

Goose is a credible default for a team that wants to own its agent stack, and a poor substitute for a well-configured commercial agent on one laptop. Its real contribution is not the coding loop, which is commodity, but the combination of a foundation-owned licence, a declarative extension allowlist and a recipe format with machine-checkable success criteria. That is the layer organisations keep rebuilding badly on top of someone else's agent.

  1. Adopt it when agents have to run on machines you do not control, or where code cannot leave the network. The binary and its documentation can both be packaged for air-gapped use.
  2. Adopt it to standardise behaviour across a fleet. The per-recipe extension allowlist and the permission mode are the mechanism, and the recipe file is the artefact you version and review.
  3. Adopt it as a second opinion. Pointing the same task at two agents with different tools and comparing the diffs is cheap when both are free to run.
  4. Do not adopt it as a developer's first agent. The install is a script and the first session needs a provider key, so the setup cost lands on whoever has the least time for it.
  5. Do not switch a working commercial agent over to it. On individual task quality the leaders are still ahead and the surrounding ecosystem is deeper.

Sources

  1. goose documentation: quickstart
  2. goose documentation: configuration files
  3. goose documentation: permission modes
  4. goose documentation: smart context management
  5. goose documentation: recipe reference
  6. goose documentation: CLI commands
  7. goose documentation: known issues
  8. goose has a new home: the Agentic AI Foundation (7 April 2026)
  9. goose repository on GitHub

Frequently asked questions

Is goose still a Block project?

No. Block donated goose to the Agentic AI Foundation at the Linux Foundation in December 2025, alongside Anthropic's Model Context Protocol and OpenAI's AGENTS.md, and the move was announced on 7 April 2026. The repository also moved from block/goose to aaif-goose/goose, so clones need their remote updated.

Can goose use my existing ChatGPT or Claude subscription?

Yes, over the Agent Client Protocol. The quickstart offers a ChatGPT subscription option that signs in with existing credentials to reach Codex models, alongside plain API keys, OpenRouter, a third-party agent router and local Ollama models.

What is the difference between recipes and sub-recipes?

A recipe is a reusable YAML unit with instructions, extensions, parameters and settings. A sub-recipe is one that another recipe invokes as a tool: it runs in a separate session with its own context and no shared memory, and cannot itself define sub-recipes. The docs still label sub-recipes experimental.

How do I stop goose from editing files it should not?

Switch the permission mode away from autonomous. Manual approval asks before every tool, smart approval auto-approves what it reads as low risk, and chat only blocks tool use entirely. Granular per-tool permissions apply in the two approval modes, and GOOSE_MAX_TURNS caps how many turns it can take unattended.

Sounds like what you need?

Tell me about your project or role – I’d love to hear from you.