Tools/AI agents

Pydantic AI review: typed Python agents with validated output

Pydantic AI 2.55 gives Python agents typed dependencies, validated output and OpenTelemetry tracing. What 2.0 changed, what Logfire costs and who should pick it.

Type
Agent framework
Pricing
MIT · free library, Logfire Team from $49 a month

··8 min read

  • Agent framework
  • Typed Python
  • Structured output
  • Dependency injection
  • OpenTelemetry
Cover art for the Pydantic AI review: a typed agent loop with tools, a validation gate and a retry path back to the model.

Key takeaways

  • Pydantic AI 2.55.0, released on 9 October 2026, is MIT-licensed and free to run. Your costs are model tokens and, if you use it, Logfire records.
  • The 2.0 line became stable on 23 June 2026 after seven betas and moved configuration onto capabilities, so pin the version and follow the upgrade path.
  • Dependencies reach your tools through a typed RunContext, and output_type validates every answer, sending a failed check back to the model before your code sees it.
  • Tracing is opt-in and follows OpenTelemetry, so spans can go to Logfire or to any OpenTelemetry backend. Durable runs need Temporal, DBOS, Prefect, Restate or AWS Lambda.
  • Pick it for typed Python services. Pick the OpenAI Agents SDK for a small OpenAI-first stack, and LangGraph when explicit graphs, checkpoints and interrupts are the product.

Listen to this article

0:000:00

Pydantic AI is the Python agent framework from the team behind Pydantic, the data validation library. Dependencies, tools and results are ordinary Python types, and every structured answer from the model is checked before your code receives it. The verdict up front – use it for typed agents inside Python services, where the team already thinks in Pydantic models. Skip it if your team builds in TypeScript, wants a visual builder as the main way to design workflows, or cannot absorb API changes between releases.

What it is

The current release is 2.55.0, published on PyPI on 9 October 2026. It is MIT-licensed, needs Python 3.11 or newer and has about 20,500 stars on GitHub. The project calls itself “how Python does AI”: agents, realtime voice, image generation and embeddings, typed end to end. The 2.0 line became stable on 23 June 2026, so older tutorials may still describe the 1.x API.

  • Typed dependencies. `deps_type` declares what an agent needs, and each run receives an instance through `deps=`.
  • Validated output. `output_type` takes a Pydantic model, a union or a list of types. The model fills it through tool calling by default.
  • About two dozen providers. A prefix such as `openai:`, `anthropic:` or `google:` selects the provider. The directory also covers Groq, Mistral, Ollama and OpenRouter, plus any OpenAI-compatible endpoint.
  • Capabilities and durability. Capabilities bundle tools, hooks, instructions and model settings into one reusable unit. Durable runs plug in through Temporal, DBOS, Prefect, Restate and AWS Lambda.

How it works

An agent run is a loop with a check at the end. The agent sends its instructions, the message history and the tool schemas to the model. When the model calls a tool, Pydantic AI runs your Python function with the typed context and returns the result to the model. When the model stops calling tools, its output is validated against `output_type`. A failed check goes back to the model as a retry, and the default budget for output retries is one.

One agent run in Pydantic AIYour code passes dependencies and input to an agent. The agent sends its instructions and tool schemas to the model. When the model calls a tool, the tool runs with the typed context and its result goes back to the model. When the model finishes, its output is validated against the output type. A failed check goes back to the model as a retry, and a passing answer is returned to your code.One agent runsame loop for every providerYour codedeps and inputAgentinstructions, toolsModelprovider prefixValidated outputoutput_typeTool callRunContext[Deps]Retrybudget 1 by defaultfails checkDependencies reach your functions, never the model.Every answer passes the validation gate before your code sees it.
Each run loops between the model and the tools, and every answer passes the validation gate before it leaves the agent.

Dependencies are passed to your functions, never to the model. A database pool or an HTTP session can sit next to the agent without appearing in a prompt. The model only sees the instructions, the tool schemas and the messages, which is what makes the boundary worth designing on purpose.

Getting started

Install with `uv add pydantic-ai` or `pip install pydantic-ai`, then set the credentials your provider expects. The example below is a support triage agent. It takes a typed dependency, calls one tool and returns a validated object.

from dataclasses import dataclass
from typing import Literal

from pydantic import BaseModel, Field
from pydantic_ai import Agent, RunContext

from myshop.orders import OrderService


@dataclass
class SupportDeps:
    customer_id: int
    orders: OrderService  # your own client, injected per run


class Triage(BaseModel):
    category: Literal['refund', 'shipping', 'other']
    risk: int = Field(ge=0, le=10, description='How urgently a person should review this')
    reply: str


support = Agent(
    'openai:gpt-6-sol',
    deps_type=SupportDeps,
    output_type=Triage,
    instructions='Triage the message. Check the order before you answer.',
)


@support.tool
async def latest_order(ctx: RunContext[SupportDeps]) -> str:
    '''Return the status of the most recent order.'''
    return await ctx.deps.orders.latest_status(ctx.deps.customer_id)


result = support.run_sync(
    'Where is my parcel?',
    deps=SupportDeps(customer_id=42, orders=OrderService()),
)
print(result.output.category, result.output.risk)

Two parts of that code do the work. The `Triage` class is both the schema sent to the model and the type your code receives. An answer outside the allowed categories or the 0 to 10 risk range is retried and, if it still fails, raises an error instead of reaching your code. The `RunContext[SupportDeps]` annotation gives the tool a typed view of your client, so the editor can check every attribute you use.

Typed dependencies and validated output

Dependencies are the part I would adopt first. `deps_type` declares the type, `RunContext[Deps]` gives tools, instructions and output validators access to `ctx.deps`, and a test can swap the real client for a fake with `agent.override(deps=...)`. The wiring stays in the constructor and the run call rather than in module-level globals, which keeps the agent easy to test.

Output is where the framework earns its keep. By default the model returns structured data through its tool-calling interface, and a union of types becomes one output tool per member. The `TextOutput` and `PromptedOutput` markers switch to plain text for models with unreliable tool calling. `ToolOutput` gives one output tool its own retry budget, so a complex type can get more attempts than a simple one.

  • `ModelRetry` lets a tool or output function reject a value and tell the model what to change.
  • `@agent.output_validator` runs your own checks after parsing, for example that an order number in the answer exists.
  • `Agent(retries={'output': N})` raises the output retry budget for the whole agent. The default is one.
  • `ToolOutput(Fruit, max_retries=2)` gives one output type its own retry count.

Testing is where the design pays back. `TestModel` calls every tool and returns a structurally valid answer, `FunctionModel` lets a test script the model’s reply, and `ALLOW_MODEL_REQUESTS=False` blocks accidental calls to real providers in CI. A unit test of the support flow then needs no API key and no network.

Tracing and durable runs

Tracing is opt-in. Call `logfire.configure()` and `logfire.instrument_pydantic_ai()` at start-up, and each run, model response and tool call becomes an OpenTelemetry span that follows the generative AI semantic conventions. The Logfire SDK can send the same data to any OpenTelemetry backend, which matters if the telemetry has to stay in a system the team already runs. For the wider case, read agent observability with OpenTelemetry.

Durable execution is the second feature to understand. The docs list eight engines. Temporal, DBOS, Prefect, Restate and AWS Lambda are co-maintained with their vendors, and Kitaru, Apache Airflow and Absurd come as external integrations. In 2.x you attach a durability capability to the agent. The README example adds `TemporalDurability()` to an agent’s capabilities inside a Temporal workflow.

Cost, hosting and data protection

As of October 2026, the library costs nothing to run. The licence covers the code, so the bills come from three places: model tokens from your provider, Logfire records if you use Logfire, and the infrastructure for a durable engine if you adopt one. The Logfire plans show the shape of the second bill.

PlanPriceIncludedWhat changes
PersonalFree10M records a month, hard-capped3 projects, 30-day retention, 1 seat and 2 read-only guests
Team$49 a month10M records, then $2 per million5 seats (up to 12), 10 guests, 5 projects, 30-day retention, spending cap
Growth$249 a month10M records, then $2 per millionUnlimited seats, guests and projects, 90-day retention, priority support and a BAA template
EnterpriseCustomBy contractCloud, Dedicated or Self-hosted; SSO, SCIM and an SLA

Logfire bills records: logs, spans and metrics. The included 10 million a month are covered by the plan credit, and above that Team and Growth charge $2 per million. A team that sends 30 million records a month pays $49 plus $40 for the extra 20 million, about $89 before any model tokens. The AI gateway adds a 5 percent markup on built-in providers, while up to three of your own provider keys pass through without markup.

Self-hosting the library is the default, because it is just code in your own environment. Logfire’s Enterprise tier adds a self-hosted option on your own Kubernetes cluster. For data protection, three flows matter. The model provider receives prompts and tool results, so its data processing terms and region come first. Logfire receives every span you export, so exclude prompts and completions at source where you do not need them. Pydantic offers a Data Processing Addendum for GDPR, a SOC 2 Type 2 report on request and a published list of subprocessors.

Region is a setting to check on every plan. The plan matrix ticks EU or US data region for each hosted plan, from Personal to Enterprise Cloud. Enterprise Dedicated offers any Google Cloud region, and self-hosting keeps the data wherever you run it. The pricing page does not say which region a new project gets by default, so check that before you send personal data. The wider data residency question for model APIs is covered in a separate article.

Where it falls short

The main risk is churn – and the changelog is open about it. The 2.0 line went through seven betas between 20 May and 10 June 2026 before the stable release on 23 June. Its breaking changes come in two groups: removals that the V1 deprecation warnings could not announce, and changes that V1 did warn about. Removed items include the Outlines integration and its extras, and `ModelProfile` changed from a dataclass to a TypedDict. A migration is a real task, not a version bump.

Minor releases are frequent too. Version 2.51.0 came out on 25 September 2026 and 2.55.0 on 9 October, so a loosely pinned project sees several changes in a fortnight. The version policy promises no intentional breaking changes in minor releases, but features in a beta module are explicitly unstable and may change in ways that break existing code. Treat any import from a beta module as a pinned dependency.

Keep an eye on the security notes as well. Release 2.52.0 fixed a CPU and memory problem in the local `web_fetch` tool, where deeply nested HTML could consume excessive resources. Provider-native web fetching was not affected. Finally – it is a Python library. Teams that build agents in TypeScript need a different framework, and durable engines add infrastructure that someone must run or buy.

ToolLicenceVersion, October 2026Strongest atTracing
Pydantic AIMIT2.55.0Typed dependencies and validated output in PythonOpt-in, through Logfire or OpenTelemetry
OpenAI Agents SDKMIT0.23.1Very few primitives: agents, handoffs, guardrails, sessionsOn by default, exported to OpenAI unless disabled
LangGraphMIT1.2.14Explicit graphs with checkpoints, interrupts and fault toleranceLangSmith, a separate platform

Verdict

Pydantic AI is my default for a Python team that already models its data with Pydantic and wants agents that are typed, testable and easy to run beside the rest of the service. It is the wrong choice for a TypeScript codebase, for a team that wants a visual graph editor as its main design tool, and for any team that cannot absorb API changes between releases. In those cases, the alternatives below fit better.

  1. Adopt Pydantic AI when your services are Python, your data is already modelled in Pydantic and you want typed tools and validated output.
  2. Pick the OpenAI Agents SDK when you are committed to OpenAI and want very few primitives. Turn tracing off or add your own processor before real customer data flows through it.
  3. Pick LangGraph when the workflow is the product: explicit state, checkpoints and named approval steps. You write more code, and you control every transition.

Sources

Frequently asked questions

How much does Pydantic AI cost?

As of October 2026, the library is MIT-licensed and free. The bills come from your model provider and, if you use Logfire, from its records: Personal is free up to 10 million records a month with ingestion paused at the cap, Team is $49 a month, Growth is $249 a month, and Enterprise is quoted.

Is the 2.x line stable enough for production?

Yes, with the usual care. Minor releases are not meant to break public APIs, but features in beta modules can change, and each major version removes what was deprecated before. Security fixes for V1 continue for at least six months after the 2.0 stable release, so plan the move.

How does it compare with the OpenAI Agents SDK and LangGraph?

The OpenAI Agents SDK is smaller and turns tracing on by default. LangGraph is built around explicit graphs with checkpoints and interrupts. Pydantic AI sits between them: typed dependencies and validated output for Python code, with a provider prefix for roughly two dozen providers.

Does Logfire receive my prompts?

Only the spans you export. Instrumentation is opt-in, the docs describe how to exclude prompts and completions from spans, and Pydantic offers a Data Processing Addendum, a SOC 2 Type 2 report on request and a list of subprocessors.

Sounds like what you need?

Tell me about your project or role – I’d love to hear from you.