Tools/AI agents
LangGraph: a low-level runtime for stateful agents
What LangGraph gives a production agent: checkpointed supersteps, interrupts and streaming, plus where the durability model stops short and what LangSmith costs.
- Type
- Agent framework
- Pricing
- Apache-2.0 · LangSmith paid
Balázs Csorba··10 min read
- LangGraph
- Agent orchestration
- Durable execution
- Human-in-the-loop
- State machines

Key takeaways
- LangGraph is an orchestration runtime, not an agent framework: it supplies a checkpointed state machine and nothing else. The 1.2 line adds per-node timeouts, error recovery, a lower-overhead channel type and a version 3 streaming API.
- Durability stops at the process boundary. A checkpointer restores state after a crash, but something outside the library has to notice the crash and re-enter the graph with the right thread_id.
- interrupt() rewinds the whole node, not the line. Any side effect before an interrupt runs again on every resume, which turns idempotency into a design constraint rather than a nicety.
- The library is MIT-licensed on GitHub and 1.2.14 was current on PyPI in October 2026; the money is in LangSmith, where Developer is free with 5k base traces a month and Plus is $39 per seat.
- The three durability modes are a real performance lever: sync commits every checkpoint before the next step starts, exit commits nothing until the run ends.
LangGraph is the part of the LangChain stack that executes the agent. Nodes and edges are declared over a shared state object, and the runtime walks the graph one superstep at a time, writing a checkpoint after each one so the run can be paused, resumed and inspected. That is the whole product, and the verdict follows from it: this is the best open-source answer to a narrow question, and teams that need the narrow question answered well should take it.
It sits below the agent and above the model. LangChain agent abstractions and the newer deepagents package are built on it; CrewAI, LlamaIndex and the OpenAI Agents SDK approach the same job from the other direction with more opinions baked in. Measured against writing the loop by hand, the contribution that matters is persistence, not orchestration.
What it is
The current release line is langgraph 1.2.x. Version 1.2.14 was on PyPI in October 2026 and requires Python 3.10 or newer. The repository credits Pregel and Apache Beam as inspirations and NetworkX as the model for the public interface, and it states plainly that the library can be used without LangChain itself.
- A graph of nodes over one state object. Nodes return partial updates, and channel reducers decide how two concurrent writes to the same key merge.
- Checkpointers, which store a state snapshot per superstep and organise runs into threads addressed by
thread_id. - Stores, a separate cross-thread key-value layer for long-term memory such as user preferences and shared reference data.
interrupt(), which suspends a node anywhere in the graph and hands control back to the caller until the graph is resumed.- Durability modes named
exit,asyncandsync, set per invocation. - Typed streaming. From 1.2,
stream_events(..., version="v3")returns separate projections for messages, values, interrupts and the final output.
Two APIs reach the same runtime. The graph API is built on StateGraph and expresses control flow as edges. The functional API expresses it as ordinary Python decorated with @entrypoint and @task. The functional version reads better in a pull request; the graph version draws as a picture, which matters more than it sounds when an operations team has to reason about what the agent will do at three in the morning.
How it works
Execution follows the Pregel model. Every node that is ready to run starts together in a superstep; when they all finish, their writes are committed as a single checkpoint and the next superstep is scheduled from the updated state. That is what makes pending writes worth caring about: if one node in a superstep throws, the writes of its successful siblings are already durable, and a resume does not re-run them. Expensive model calls inside a fan-out are paid for once.
Concurrency is where the sharp edges live. Two nodes writing the same state key in the same superstep need a reducer. LangGraph ships add and a last-value default, and the rest is written by hand. Getting this wrong does not raise anything: one of the two writes is silently dropped, and the symptom shows up much later as a decision the agent cannot explain.
State is also the storage bill. Every superstep rewrites the whole state object, so a key that accumulates retrieved documents turns each step into a full payload write. Prune the state at the node boundary and keep the durable artefacts outside the graph.
Getting started
The smallest graph worth writing is an approval flow, because that is the case the runtime is actually for. This one pauses for a human, resumes on the same thread, and keeps the side effect after the interrupt so it executes exactly once.
from typing import Literal, TypedDict
from langgraph.checkpoint.postgres import PostgresSaver
from langgraph.graph import END, START, StateGraph
from langgraph.types import Command, interrupt
class State(TypedDict):
request: str
decision: str | None
def ask(state: State) -> Command[Literal["send", "cancel"]]:
if interrupt({"question": "Send this?", "details": state["request"]}):
return Command(goto="send")
return Command(goto="cancel")
builder = StateGraph(State)
builder.add_node("ask", ask)
builder.add_node("send", lambda s: {"decision": "sent"})
builder.add_node("cancel", lambda s: {"decision": "cancelled"})
builder.add_edge(START, "ask")
builder.add_edge("send", END)
builder.add_edge("cancel", END)
with PostgresSaver.from_conn_string("postgresql://…") as saver:
saver.setup()
graph = builder.compile(checkpointer=saver)
config = {"configurable": {"thread_id": "req-42"}}
print(graph.invoke({"request": "refund 8891"}, config, durability="sync")["__interrupt__"])
print(graph.invoke(Command(resume=True), config, durability="sync")["decision"])Two details in that snippet matter more than the rest. PostgresSaver.setup() creates the checkpoint tables once, at deploy time, not once per process. And durability="sync" is the right choice for an approval flow: the write completes before the next step, so a crash in the two seconds after the decision cannot lose the decision itself.
Persistence and durability
Persistence is the reason to adopt the framework and also where the marketing language gets loose. A checkpointer writes a snapshot. Durable execution means the run continues. LangGraph does the first; the second is left to whoever operates the process.
exit— nothing is written until the run completes, fails or interrupts. Fastest, and a process crash loses the run.async— writes run while the next step executes. The default trade: good latency, and a small window in which a crash loses state.sync— every checkpoint is committed before the next step starts. Highest durability, with the write on the critical path.
| Package | Backend | Where it fits |
|---|---|---|
langgraph-checkpoint | In memory | Tests and experiments; ships with langgraph |
langgraph-checkpoint-sqlite | SQLite | Local workflows and single-process apps |
langgraph-checkpoint-postgres | PostgreSQL | Production; also what LangSmith Deployment runs on |
langgraph-checkpoint-mongodb | MongoDB | Teams already standardised on MongoDB |
langchain-azure-cosmosdb | Cosmos DB | Azure shops, with Entra ID authentication |
Checkpoints grow without bound. The persistence documentation says so directly and suggests a scheduled job that deletes checkpoints older than a retention window. Teams that skip this discover it through a database that has quietly become the largest thing in the stack, and that is a bad afternoon.
Interrupts and the replay rules
interrupt() is the most-used feature and the most misread one. It does not pause at a line. It raises an exception, unwinds to the runtime, checkpoints the state and waits indefinitely. When the graph resumes, the runtime restarts the entire node from the top and matches resume values to interrupt calls strictly by index. Every production bug in this area comes from ignoring one of those two sentences.
- Never wrap an
interrupt()call in a baretry/except. The pause is a thrown exception and a broad handler swallows it, so the graph never pauses at all. - Do not conditionally skip or reorder interrupts inside a node. Matching is index-based, so a changed call order consumes the wrong resume value without any error.
- Make every side effect before an interrupt idempotent, or move it after the pause, or split it into its own node. The documentation is explicit that a record created before the interrupt is created again on each resume.
- Avoid
while Trueloops around an interrupt in a single node. Every resume replays the earlier iterations, so work inside the loop grows exponentially.
Where it falls short
The weaknesses come first, because they are what decides whether the framework fits. A run lives in one process. There is no supervisor, no task queue and no worker pool in the open-source library, so if that process dies the run is dead until a system outside LangGraph notices and re-enters it. Human review has the same shape: interrupt() halts the run, and building the thing that notices an approval arrived and wakes the right thread is now your problem.
| LangGraph | CrewAI | LlamaIndex | |
|---|---|---|---|
| Control model | Explicit graph or functional API | Roles and tasks | Composable pipelines and indices |
| Persistence | Checkpoints per superstep, you run the process | Memory and knowledge abstractions | Checkpointing inside workflow and index nodes |
| Strongest at | Long, resumable, auditable runs | Fast multi-agent prototypes | Retrieval-heavy applications |
| What it leaves you | Prompts, tool loop, retries, supervision | Fine control of the run itself | Orchestration tied to retrieval |
The second weakness is ergonomics. LangGraph abstracts nothing about prompts or architecture, which is a virtue when the agent is the product and an obstacle when it is not. The framework will not tell you how to structure a prompt, when to stop calling tools, or how many retries a step deserves. Most teams spend the first weeks of a LangGraph project rediscovering the tool loop, which is exactly what a higher-level abstraction would have handed them.
The third is lock-in, and it is milder than it usually is claimed to be. The runtime is MIT-licensed, runs in your process and writes to your database, so there is no data held hostage. The real dependency appears once graphs are deployed through LangSmith: the deployment, the assistants API and the cron scheduling are LangChain surfaces, and moving off them later is real work.
What it costs
LangGraph is free. LangSmith is where the money goes, and it is priced per seat with metered usage on top: Developer is free for one seat with 5,000 base traces a month, Plus is $39 per seat per month with 10,000 base traces and access to deployment, Engine and sandboxes, and Enterprise is priced on request with hybrid or fully self-hosted deployment.
- Usage is metered in LangChain Standard Units at $1 each, and a serverless deployment is billed on runtime compute, runtime memory, database compute and database memory, plus the time the database is live.
- Trace retention is 14 days for a base trace and 180 days for an extended trace, which is billed separately.
- LangSmith states that it does not train models on customer data, and offers a self-hosted data plane on the Enterprise tier for teams whose own controls require it.
The trace allowance is the number to watch. One agent run is one trace, and a run that calls five tools across ten nodes produces a graph of spans inside it. Tracing every production request exhausts 5,000 or 10,000 traces surprisingly fast, which is the moment the bill stops being a rounding error. Sampling by environment is the cheapest mitigation.
Verdict
LangGraph is a good answer to one question: how do I keep a long agent run alive across a crash, a deploy or a human approval. Inside that boundary the work is careful — the checkpoint format is documented, pending writes are a genuinely good idea, and the interrupt semantics are spelled out plainly enough to design around. Outside it, the library is a runtime with no opinions, and everything it declines to decide becomes the reader’s work at three in the morning.
- Adopt it when a run must survive a restart: an approval queue, a research job that runs for hours, an agent that waits for a person.
- Adopt it when the run has to be auditable. Nodes and edges are the cheapest way to show a non-engineer exactly what the agent will do.
- Adopt it when mixing deterministic and model-driven steps matters and the boundary between them has to be exact and testable.
- Skip it for a tool-calling loop that finishes in three steps. Twenty lines of Python are cheaper to own and faster to debug.
- Do not treat it as the durability layer. If a silently dead run means lost orders, add a supervisor or move the graph onto Temporal, whose LangGraph plugin went to public preview in July 2026.
- Do not adopt it without reading the interrupt rules once. The replay semantics are the part that produces duplicate side effects in production.
LangGraph is a simple, efficient way to express an agent: the graph model is clear, the ecosystem is rich, and prototypes come together fast. But, it is not a complete production story. Temporal, on its LangGraph plugin, July 2026
Sources
Frequently asked questions
Is LangGraph free to use in production?
The library is MIT-licensed, and the separate checkpointer packages are MIT too. LangSmith is optional: the Developer tier is free for one seat with 5,000 base traces a month, Plus is $39 per seat per month with 10,000 traces. Nothing in LangGraph requires a LangSmith account, though debugging a multi-node run without traces is close to guesswork.
Do I need LangGraph for a simple agent loop?
Probably not. A tool-calling loop is a while loop around a model call and costs about twenty lines. LangGraph earns its place when a run has to survive a restart, wait for a human, or resume from a named step, because a plain loop has no way to represent any of that.
What does it mean that checkpointing is not durable execution?
A checkpointer writes graph state at every superstep; durable execution means the run itself continues. Nothing in the library restarts a dead run or stops two processes from resuming the same thread at once. Temporal shipped a LangGraph plugin in public preview in July 2026 that runs the graph as a Temporal workflow to close exactly that gap.
What does checkpointing actually cost?
One database write per node per superstep, carrying the state payload. The write latency depends on the durability mode, but storage is the bigger problem: the docs recommend a scheduled job that deletes checkpoints older than a retention window, because they grow without bound.