Tools/AI agents
OpenAI Agents SDK: a small agent runtime with sharp edges
A review of the OpenAI Agents SDK: the runner loop, tracing, guardrails and approvals, plus what the release churn and the Responses-only features cost.
- Type
- Agent framework
- Pricing
- MIT · API pay per token
Balázs Csorba··10 min read
- Agent runtime
- Tracing
- Guardrails
- MCP
- Python

Key takeaways
- The Python package is MIT licensed, needs Python 3.10 or newer, and stood at version 0.23.1 on 2 October 2026 after 123 PyPI releases since March 2025.
- The Runner caps a run at max_turns=10 by default and raises MaxTurnsExceeded, which is a sensible default most teams should keep.
- Tracing is enabled by default and carries model and tool inputs and outputs, so production runs should set trace_include_sensitive_data=False.
- Guardrails run beside the agent by default, so tokens are already spent when a tripwire fires; run_in_parallel=False is the setting for cost-sensitive paths.
- Computer use, hosted tool search and programmatic tool calling are rejected on Chat Completions models and on non-Responses backends, which makes the provider-agnostic claim thinner than it reads.
OpenAI Agents SDK is the agent runtime OpenAI ships for Python and TypeScript: a small set of primitives, a turn loop, and tracing that is switched on before anyone asks for it. It is a good default for a team standardising on OpenAI models that would rather the loop belonged to a library than to hand-written asyncio. It is the wrong tool the moment the workflow turns into a state machine.
It sits between the raw Responses API and a full orchestration framework. The Responses API is the model interface. The SDK adds a Runner that owns turns, tool dispatch, guardrails, handoffs and sessions. Graph runtimes such as LangGraph sit above both and encode the workflow explicitly. The SDK documentation draws the line itself: use the Responses API directly when the intention is to own the loop, tool dispatch and state handling.
What it is
The design is deliberately thin. An Agent is instructions plus a model plus tools. Delegation has exactly two shapes: Agent.as_tool() for a manager that keeps the conversation, and handoff() for a specialist that takes it over. Guardrails are ordinary functions that return a tripwire flag. Everything else, sessions, MCP, sandbox clients, realtime and voice, is a module around that core rather than a new abstraction to learn first.
- Package:
openai-agentson PyPI, MIT licensed, requires Python 3.10 or newer. - Version: 0.23.1 on 2 October 2026, the 123rd release since 0.0.1 shipped on 4 March 2025; 77 of those releases landed in 2026 alone.
- Models: the OpenAI Responses API by default, Chat Completions as an explicit alternative, LiteLLM and AnyLLM adapters for other providers.
- Tools: plain Python functions behind the @function_tool decorator, plus hosted tools, computer use, shell and apply-patch.
- MCP: stdio, Streamable HTTP, the deprecated SSE transport and hosted MCP servers, each with tool filters and approval policies.
- Memory: SQLite, SQLAlchemy, Redis, MongoDB, Dapr, encrypted and OpenAI Conversations-backed sessions.
- Durable execution: not in the box; the documentation routes long-running runs to Temporal, DBOS, Dapr or Restate.
How the loop works
Runner.run() is a loop, not a function call. It sends the current input to the model, then does one of three things with what comes back: treats text of the expected output type with no tool calls as the final output, switches to another agent on a handoff, or executes the requested tools, appends their results and goes round again.
That definition of final output is worth reading twice, because it is what makes the loop a loop. A run finishes when the model produces text of the requested type and asks for nothing. Anything else keeps it alive, which is why max_turns is the first setting to decide deliberately rather than inherit. Passing max_turns=None removes the limit entirely.
Getting started
pip install openai-agents is the whole setup, and the SDK reads OPENAI_API_KEY when it first creates a client. A minimal agent with a tool and a typed answer is about twenty lines:
from pydantic import BaseModel
from agents import Agent, Runner, function_tool
@function_tool
def order_status(order_id: str) -> str:
"""Look up the fulfilment status of an order."""
return STATUS.get(order_id, "unknown")
class Reply(BaseModel):
answer: str
order_id: str
agent = Agent(
name="Order assistant",
instructions="Answer with the status of the order the customer names.",
tools=[order_status],
output_type=Reply,
)
result = Runner.run_sync(agent, "Where is order A-1024?", max_turns=6)
print(result.final_output)
print(result.context_wrapper.usage.total_tokens)The decorator derives the JSON schema from the signature and the docstring, so the model sees exactly as much as that docstring states. output_type turns the final message into a validated Pydantic model instead of a string. usage is aggregated across every model call in the run, including the ones that produced tool calls and handoffs.
Guardrails and approvals
Guardrails are the part that gets misunderstood most often, because they do not all fire at the same point. Input guardrails run only for the first agent in a chain, output guardrails only for the agent that produces the final output, and neither of them looks at the delegated work in between.
- Set run_in_parallel=False on an input guardrail to block the agent before it starts. The default runs the guardrail beside the agent, which lowers latency but means tokens are already spent when a tripwire fires.
- Tool guardrails wrap individual function tools and local MCP tools, and are the only kind that sees every call in a multi-agent chain.
- A tripwire raises InputGuardrailTripwireTriggered or OutputGuardrailTripwireTriggered. A guardrail function that raises is treated as an unknown verdict and the runner persists the completed turn before surfacing the error.
- needs_approval on a tool, on Agent.as_tool(), on ShellTool or on ApplyPatchTool pauses the run instead; the pending calls appear in result.interruptions.
- Callable approval rules fail closed. If the arguments are missing, malformed or not a JSON object, the call requires manual approval rather than being waved through.
Approval state is serialisable through RunState, so a paused run can sit in a queue and resume in another process. The documentation is explicit that RunState.from_json() authenticates nothing: a snapshot in untrusted hands is a set of instructions the server will execute, so it belongs in server-side storage with the reviewer authenticated by the application. That is the same approval pattern as human-in-the-loop review, with a state file instead of a socket.
Tracing and cost control
Tracing is the strongest reason to pick this SDK and the least controlled part of it. Spans cover the runner, each task and turn, each agent, each generation, each function call, guardrails, handoffs and audio. The default BatchTraceProcessor exports in the background every few seconds, which means a worker can finish a job and exit before the dashboard shows the run.
- Spans emitted by default: runner, task, turn, agent, generation, function, guardrail, handoff, transcription and speech.
- Turn it off globally with OPENAI_AGENTS_DISABLE_TRACING=1 or set_tracing_disabled(True), or for one run with RunConfig(tracing_disabled=True).
- For a delivery guarantee, call flush_traces() after the trace context closes. Disabling tracing does not discard spans that a processor has already buffered.
- add_trace_processor() adds a destination and leaves the OpenAI exporter registered. set_trace_processors() replaces the default and needs its own BatchTraceProcessor with an exporter.
- The documentation lists roughly 27 external processors, among them Langfuse, MLflow, Arize Phoenix, LangSmith, Braintrust, Datadog and Pydantic Logfire.
Cost control is arithmetic rather than configuration. result.context_wrapper.usage carries requests, input_tokens, output_tokens, total_tokens, per-request entries and cached and reasoning token detail; the compaction request a Responses session issues is added to the same totals. The number that matters is cost per successful run, and the lever with the most leverage is still max_turns, because a loop that needs twelve turns to fail will spend twelve turns every single time.
Where it shingles
Start with the weaknesses. Durable execution is not in the box: a run that must survive a process restart, wait hours for an approval or resume in a new container needs Temporal, DBOS, Dapr or Restate alongside it. The Python sandbox and harness work landed in 0.14.0 in April 2026, and TypeScript support was still described as future work at the time of writing. Tracing ships payloads to OpenAI by default and is simply unavailable under ZDR. And the provider-agnostic claim is thinner than it sounds: computer use, hosted tool search and programmatic tool calling are rejected on Chat Completions models and on non-Responses backends.
Then there is churn. The project shipped 77 releases in the first nine months of 2026, and 0.21.0 raised the floor to openai>=3.0.0,<4, which moved the default provider onto HTTPX2 and broke applications that passed a legacy httpx client. A version still prefixed 0.x at that cadence is an argument for pinning and for reading the changelog rather than the release headline.
| Option | Control model | Durable execution | Latest release, October 2026 |
|---|---|---|---|
| OpenAI Agents SDK | Model-directed loop, handoffs, agents as tools | External: Temporal, DBOS, Dapr, Restate | 0.23.1 |
| LangGraph | Explicit graph with checkpoints | Built in | 1.2.14 |
| Pydantic AI | Code orchestration, typed agents | External: Temporal, DBOS, Prefect, Restate | 2.54.0 |
| Google ADK | Model-directed agents plus workflow agents | Built in, ADK 2.0 workflow runtime | 2.11.0 |
Read the fourth column before the second. LangGraph at 1.x and ADK at 2.x sit under a compatibility promise; 0.23.1 means OpenAI can move a constructor signature in a patch release, which it has already done once with the openai client floor. The control-model column matters less than it looks: a model-directed loop is faster to build, and a graph is easier to reason about only for as long as the graph stays small.
Verdict
The SDK wins on the things that are hard to retrofit. Tracing that exists before the observability code is written, approval interrupts that serialise cleanly, and MCP across four transports without a hand-rolled client are all real engineering time saved. It loses on the things that cannot be added later without a rewrite: durable execution, and an explicit control flow that makes a workflow auditable.
- Adopt it when the team is standardising on OpenAI models and the workflow is a tool loop with guardrails and approvals. That is the case it was designed for.
- Adopt it when tracing is non-negotiable, because the span model and the dashboard are already wired to the runtime instead of bolted on afterwards.
- Adopt it for MCP-heavy agents, where four transports and per-server tool filters save more code than the framework costs.
- Do not adopt it as the only runtime if runs must survive a restart, wait for a human for hours, or resume in a fresh container. Pair it with Temporal or DBOS from the first day.
- Do not adopt it for a workflow whose next step is a business rule. If the sequence is known, a graph runtime or plain code says it more honestly than a model that might not follow it.
- Skip it entirely when the whole job is one model call returning one response. The Responses API plus Pydantic is smaller and cheaper.
None of that is a knock-on the package. It is a small, readable, MIT-licensed library that does a narrow job properly and says so in its own documentation. The failure mode to avoid is not picking it; it is picking it because the tracing is free and then discovering a year later that the workflow logic lives in prompts nobody can diff. Keep durable execution and the business rules outside the runner.
Enough features to be worth using, but few enough primitives to make it quick to learn. — the two design principles stated in the OpenAI Agents SDK documentation.
Sources
- OpenAI Agents SDK documentation: Intro, and Agents SDK or Responses API
- OpenAI Agents SDK documentation: Running agents
- OpenAI Agents SDK documentation: Guardrails
- OpenAI Agents SDK documentation: Human-in-the-loop
- OpenAI Agents SDK documentation: Tracing
- OpenAI Agents SDK documentation: Configuration
- openai-agents 0.23.1 on PyPI, release history and licence
- OpenAI: The next evolution of the Agents SDK (15 April 2026)
- OpenAI API documentation: Agents, comparison of the three runtimes
- Arize: AI agent frameworks compared (1 October 2026)
Frequently asked questions
Is the OpenAI Agents SDK free?
The SDK itself is MIT licensed and installs with pip install openai-agents. The money goes to the API: the models the loop calls are billed per token, and OpenAI states that the harness and sandbox capabilities from April 2026 use standard API pricing based on tokens and tool use.
Can the OpenAI Agents SDK run non-OpenAI models?
It can. The package ships LiteLLM and AnyLLM adapters, reads OPENAI_BASE_URL, and the project README claims support for 100 or more models. Several capabilities are Responses-only and are rejected on Chat Completions models and on non-Responses backends, so test a complete agent run rather than a single model call.
What is the difference between the Agents SDK and the Responses API?
The Responses API is the model interface; the SDK adds a runtime around it that owns turns, tool dispatch, guardrails, handoffs and sessions. If the job is one call that returns one response, the SDK adds machinery nothing uses, and the Responses API is the smaller choice.
How does human approval of tool calls work?
Set needs_approval on a function tool, on Agent.as_tool(), on ShellTool or on ApplyPatchTool, and the run pauses with ToolApprovalItem entries in result.interruptions. Convert the result with to_state(), call state.approve() or state.reject(), and resume with Runner.run(agent, state). Callable approval rules fail closed when the arguments cannot be parsed.