> Where approval gates belong in an AI agent, how to avoid rubber-stamping, and how interrupt and resume work in LangGraph and the OpenAI and Claude agent SDKs.
>
> Web page: https://balazscsorba.com/blog/human-in-the-loop-ai-agents · Language: English · Also available in: [Deutsch](https://balazscsorba.com/de/blog/human-in-the-loop-ai-agents.md) · [Magyar](https://balazscsorba.com/hu/blog/human-in-the-loop-ai-agents.md)
> Author: Balázs Csorba · Published: 2026-10-02 · Keywords: human in the loop AI agents, human in the loop KI-Agent, Mensch im Loop KI, AI agent approval gates, agent approval fatigue, LangGraph interrupt human in the loop, OpenAI Agents SDK needs_approval, Claude Agent SDK canUseTool, AI agent audit trail, risk tiers for AI agent actions

[Blog](https://balazscsorba.com/blog)/AI agents

# Human in the loop for AI agents: where to put approval gates

Where approval gates belong in an AI agent, how to avoid rubber-stamping, and how interrupt and resume work in LangGraph and the OpenAI and Claude agent SDKs.

[Balázs Csorba](https://balazscsorba.com/about)·October 2, 2026·13 min read

-   Human in the loop
-   AI agents
-   Approval gates
-   Agent safety

![Diagram: an agent proposes an action, a risk gate sends it to automatic execution, to a human approval, or to a block, and every decision lands in an audit log.](https://balazscsorba.com/images/blog/human-in-the-loop-ai-agents/cover.webp?v=66ffffa4d3)

## Key takeaways

-   Gate actions by risk, not by tool: reversibility, blast radius, data sensitivity and who sees the effect decide whether an action runs, asks a human or is blocked.
-   Per-action approval stops working under volume. Anthropic reports that users approve 93 percent of Claude Code permission prompts and that, in one experiment, human review caught only 13.6 percent of a disguised dangerous command.
-   Interrupt and resume is now a standard primitive: LangGraph interrupt with Command(resume), needs\_approval and RunState in the OpenAI Agents SDK, canUseTool and the PreToolUse defer decision in the Claude Agent SDK.
-   An approval is only worth something if it is bound to the exact action, made by an authenticated person, recorded, and able to expire. Everything else is a ritual.
-   Always-on agents such as OpenAI dots move the human from the start of the chain to the end, so the design shifts from asking more often to asking better, with rules, a reviewer model and sampled audits.

On this page

1.  [Why per-action approval fails at scale](https://balazscsorba.com/#why-approval-fails)
2.  [Where to put the gates: risk, reversibility, blast radius](https://balazscsorba.com/#where-to-gate)
3.  [Approval UX that does not train people to click yes](https://balazscsorba.com/#approval-ux)
4.  [Interrupt and resume in the main frameworks](https://balazscsorba.com/#interrupt-resume)
5.  [Audit trails and escalation](https://balazscsorba.com/#audit-escalation)
6.  [What always-on agents change](https://balazscsorba.com/#always-on-agents)
7.  [A checklist for your next agent](https://balazscsorba.com/#checklist)
8.  [The bigger picture](https://balazscsorba.com/#the-bigger-picture)
9.  [Sources](https://balazscsorba.com/#sources)

Every team that ships an agent hits the same question within a week: when does it have to ask? Ask too rarely and one bad tool call deletes a table, sends the wrong email or refunds the wrong customer. Ask too often and people click "allow" without reading, which is worse than no gate, because the system now looks supervised when it is not.

"Human in the loop" is usually treated as a checkbox: add a confirmation dialog before dangerous tools. That is the start of the design, not the end. This article covers the decisions that matter: where the gates go, how to build approvals people actually read, which interrupt and resume primitives the main agent frameworks offer today, what an audit trail has to contain, and how always-on agents like [OpenAI dots](https://balazscsorba.com/blog/openai-dots-always-on-agents-impact) change the picture. If you want the mechanics of the loop itself first, start with [the agent loop explained](https://balazscsorba.com/blog/agent-loop-explained).

## Why per-action approval fails at scale

The uncomfortable evidence comes from the people who run the biggest approval prompt in the industry. In its [engineering write-up on auto mode](https://anthropic.com/engineering/claude-code-auto-mode), Anthropic states that Claude Code users approve 93 percent of permission prompts, and describes the consequence as approval fatigue, where people stop paying close attention to what they approve.

A reported experiment makes the point sharper. According to [DevOps.com](https://devops.com/anthropic-makes-claude-codes-auto-mode-the-default-betting-automation-beats-manual-review/), Anthropic embedded a genuinely dangerous command in permission prompts for 1,053 testers: human reviewers caught it 13.6 percent of the time, an automatic classifier caught it 89 percent of the time, and after more than 50 prior prompts the human catch rate fell to about 5 percent. This is a vendor experiment with a deliberately disguised command, so I read it as a direction rather than a constant. The direction is clear anyway: a human who has seen fifty harmless requests is a poor detector for the fifty-first.

The same engineering post is honest about the other side. Its two-stage classifier still showed a 17 percent false-negative rate on real overeager actions in a small sample of 52, which Anthropic calls the honest number. Neither a person nor a model is a reliable single gate. The design goal is therefore not "a human approves everything", it is **put scarce human attention where it changes the outcome, and back it with controls that do not get tired**.

## Where to put the gates: risk, reversibility, blast radius

Gating by tool name ("ask before Bash") is too coarse and too easy to dilute. I gate by four properties of the _action_, in this order of importance:

-   **Reversibility.** Can the effect be undone cheaply and completely? A draft, a branch or a row in a staging table can. A sent email, a payment, a deleted bucket or a changed permission cannot.
-   **Blast radius.** How many records, customers or systems does one call touch? "Update one ticket" and "update every ticket matching a filter" are the same tool with a thousandfold difference in risk.
-   **Data and visibility.** Does the action expose personal, financial or confidential data, or is it visible outside your organisation?
-   **Authority.** Does it use credentials the human did not grant for this task, or change who can do what?

These properties map to four tiers. The right-hand columns matter as much as the left: a tier is only real if it names the control and the evidence it leaves behind.

Tier

Typical actions

Control

Evidence

**0: observe**

Read inside the agent scope, search, summarise, draft in a scratch area

Run automatically, least-privilege read access

Log of calls

**1: reversible change**

Commit to a branch, edit a draft, create a ticket, write to a staging store

Run automatically with undo; reviewer model or rules on top

Log plus before and after state

**2: visible or costly**

Email to a customer, post in a shared channel, bulk update within a limit, spend below a budget

Human approval with the real effect shown; batch similar steps

Approver, exact payload, timestamp

**3: irreversible or privileged**

Payments, deletions, production deploys, permission changes, bulk export of personal data

Hard block or mandatory handoff; two people for the worst cases; the agent cannot lower this tier

Approver, payload, reason, second approver

Two rules keep the matrix honest. First, **escalate by the worst case of the arguments, not the tool**: a refund tool called with an amount over the limit is tier 3, under it tier 2. Second, **the agent never classifies its own action**. The tier comes from deterministic rules or from a separate component outside the agent's reach. That is exactly the structure OpenAI describes for dots, where a separate review system outside the environment the agent can change decides whether a step runs. Anthropic's permission order in the Claude Agent SDK is built the same way: [hooks, then deny rules, then ask rules, then the permission mode, then allow rules, then your callback](https://code.claude.com/docs/en/agent-sdk/permissions).

The tier is decided outside the agent. The human sees only the cases that need a human.

## Approval UX that does not train people to click yes

When a gate does fire, the interface decides whether it works. These are the design rules I apply, and they all follow from one idea: the reviewer should be able to judge the _effect_ in a few seconds.

-   **Show the effect, not the intent.** "The agent wants to run send\_email" is useless. "Email to ceo@client.com, 340 words, attaches contract.pdf, cc'ing 12 people" is reviewable. For code and data, show a diff; for money, show amount, currency and recipient.
-   **Make the dangerous part loud.** Highlight what is unusual: a new recipient, an amount above the median, a wildcard in a path, an external domain.
-   **Batch what belongs together.** Ten similar tier-2 steps should be one decision with a visible list, not ten dialogs. Fatigue grows with the count of interruptions.
-   **Offer edit, not just yes or no.** The Claude Agent SDK lets your callback return modified input, and LangChain's human-in-the-loop middleware has an edit decision next to approve and reject. A reviewer who can fix a parameter will not reject the whole plan.
-   **Make rejection carry a reason.** A deny message goes back to the model, which can then adjust. In the Claude Agent SDK, Claude sees the message of a denial; in the OpenAI Agents SDK you can set a rejection message per call.
-   **Be careful with "always allow".** Remembering a decision removes future friction and future scrutiny together. If you offer it, scope it narrowly (this command, this path), let it expire and list active rules somewhere people can review.
-   **Fail closed.** If nobody answers, the answer is no. Never let a timeout approve.

Then measure. My rule of thumb, which is an experience-based heuristic and not a published threshold: if a gate is approved more than about 95 percent of the time and the median decision takes a couple of seconds, it is theatre. Either the action belongs in tier 1, or the prompt does not give people what they need to judge it. Track approval rate, decision time and the share of approvals later reversed, per action type. This is the same attention problem that makes [agent-written pull requests a review bottleneck](https://balazscsorba.com/blog/ai-generated-pr-review-bottleneck), and the fix is the same: fewer, better-prepared decisions.

## Interrupt and resume in the main frameworks

Technically, a gate is a pause: the agent proposes an action, the runtime stops, state is saved, and a later call resumes with a decision. All three major stacks now have a first-class primitive for it. The API surface differs, the failure modes are similar.

Framework

Pause

Resume

Persistence and gotchas

**LangGraph**

Call `interrupt(payload)` in a node; the payload must be JSON-serialisable

Invoke again with `Command(resume=value)` on the same `thread_id`; the value becomes the return of `interrupt`

Needs a checkpointer. The node restarts from its beginning, so side effects before the interrupt must be idempotent. Never wrap `interrupt` in try/except. Matching is index-based, keep call order stable

**LangChain agents**

`HumanInTheLoopMiddleware` with `interrupt_on` per tool

`Command(resume={"decisions": [...]})` with approve, edit, reject (or respond)

Needs a checkpointer and a `thread_id`; decisions must match the order of the paused actions

**OpenAI Agents SDK**

A tool sets `needs_approval` (true or an async function of the arguments); the run ends with pending `interruptions`

`state.approve(item)` or `state.reject(item, rejection_message=...)`, then run again with the `RunState`

The state serialises with `to_json` and `from_json`. Only deserialise from trusted storage. `always_approve` makes a decision sticky

**Claude Agent SDK**

`canUseTool` fires for calls that no hook, rule or mode has settled; for slow reviewers a `PreToolUse` hook can return `defer`

Return `allow` (optionally with `updatedInput`) or `deny` with a message; a deferred call resumes with `--resume` and the hook fires again

Auto-approved tools never reach `canUseTool`; use a `PreToolUse` hook for checks that must see every call. In `dontAsk` mode prompts become denials

A few details from the official documentation are worth knowing before you build on them. LangGraph's [interrupts guide](https://docs.langchain.com/oss/python/langgraph/interrupts) warns that a resumed node re-runs from the top, so an email sent before the interrupt would be sent twice. The OpenAI Agents SDK [guide](https://github.com/openai/openai-agents-python/blob/main/docs/human_in_the_loop.md) shows that `needs_approval` can be an async function that decides per call from the parameters, and says that for long-lived approvals the server should authenticate the reviewer, authorise against the stored run, validate decision IDs against server-owned state and apply decisions atomically to prevent replay. The Claude Agent SDK [documentation](https://code.claude.com/docs/en/agent-sdk/user-input) notes that the callback can stay pending indefinitely, and recommends the `defer` decision when a person might take longer than your process can stay alive.

```
from langgraph.types import interrupt, Command

def refund_node(state):
    # runs again from the top after resume: keep everything above idempotent
    decision = interrupt({
        "action": "refund", "amount": state["amount"],
        "customer": state["customer_id"], "tier": 3,
    })
    return Command(goto="execute" if decision["approved"] else "cancel")

# later, possibly from another process, same thread_id
graph.stream_events(Command(resume={"approved": True}),
                    config={"configurable": {"thread_id": "case-4711"}}, version="v3")
```

Whatever the framework, I add four properties on top. **Bind the approval to the exact action**: store a hash of tool name and arguments with the decision and re-check it at execution, so a plan that changed after approval is not covered. **Authenticate the approver** from your session, never from the request body. **Expire approvals** after minutes or hours, not days. **Make the executing step idempotent** with an idempotency key, because resume, retry and double-click all exist. For tool security more broadly, my [MCP server security checklist](https://balazscsorba.com/blog/mcp-server-security-checklist) covers the other half of the problem.

## Audit trails and escalation

An approval that leaves no record cannot be reviewed, disputed or learned from. For each gated action I log a small, boring set of facts: the run and thread identifiers, the proposed action with full arguments, the tier and the rule that assigned it, who decided (a person, a rule or a reviewer model), the decision and any edit, the timestamp and latency, and the result of the execution. Store the arguments as they were shown to the reviewer. If the UI renders a summary, keep the summary too, because a dispute is often about what the person _saw_.

For regulated work this is not optional. Article 14 of the EU AI Act, which applies to high-risk systems, [asks for oversight](https://artificialintelligenceact.eu/article/14/) that lets assigned people understand the system's limits, stay aware of automation bias, decide not to use or to override the output, and intervene or stop the system. Whether your agent is high-risk depends on the use case, so check that before assuming either way. My [EU AI Act checklist for developers](https://balazscsorba.com/blog/eu-ai-act-article-50-developer-checklist) covers the transparency side.

Escalation is the part teams forget. Decide up front who is asked, how long they have and what happens next. A workable ladder is: the requesting user first, then a named role (the process owner), then a second approver for tier 3, and a default of reject when nobody answers. Add a **kill switch** that is independent of the agent: a flag that stops new runs and cancels pending approvals. Route notifications through a channel people already watch. The Claude Agent SDK, for example, has a PermissionRequest hook meant for sending a Slack or e-mail notification when an agent is waiting.

## What always-on agents change

Everything above assumed a person sitting in front of a chat. Always-on agents break that assumption. A dot or a scheduled coding agent works for hours, runs while you sleep and produces approvals at machine pace. As I wrote in the [dots impact analysis](https://balazscsorba.com/blog/openai-dots-always-on-agents-impact), the human moves from the start of the chain to the end, and attention becomes the bottleneck.

OpenAI's published design for dots is a useful reference because it is a risk-tiered gate. Background research runs on read-only tools, a separate Auto-review system checks consequential steps against your instructions, your Custom Rules and safety requirements, and Custom Rules can allow, require approval for or block actions but cannot remove a mandatory floor: changing a password or moving money between accounts always returns to the person. That is tiers 0, 2 and 3 in a product.

Three consequences for your own design follow. First, **rules replace most prompts**. A rule set that says "refunds under X run, over X ask, anything touching bank details is blocked" removes thousands of low-value interruptions and turns the remaining ones into events worth reading. Second, **a reviewer model is a tool, not an oracle**. It scales and it does not get tired, but Anthropic's own numbers show a non-zero miss rate, and OpenAI's Auto-review is the vendor's model supervising the vendor's agent. Keep deterministic tier-3 rules outside any model, and sample the auto-approved actions for human audit. Third, **gates control actions, not what the agent reads**. A read-only dot still sees the whole customer record, which is why least-privilege connections and masked data belong in the design as much as approvals do. I go deeper on that pattern in [prompt injection and the lethal trifecta](https://balazscsorba.com/blog/prompt-injection-lethal-trifecta-patterns).

**Ask better, not more**

When the agent runs around the clock, the number of questions is a cost you pay in human attention. Move volume out of the human queue with rules and a reviewer model, and spend the human on tier 2 and 3 decisions presented with their real effect.

## A checklist for your next agent

This is the list I would run through before putting an agent with write access in front of real data:

1.  List every tool and, for each, the worst-case arguments. Assign a tier from reversibility, blast radius, data and authority.
2.  Put the tier decision outside the agent: deterministic rules first, a reviewer model second, never the agent's own judgement.
3.  Block tier 3 by default and allow it only through a named handoff with an authenticated approver.
4.  Build the approval screen around the effect: payload, diff, recipient, amount, with unusual parts highlighted.
5.  Batch similar tier-2 steps, support edit as well as approve and reject, and return rejection reasons to the agent.
6.  Pause with a checkpointer or serialised state, make everything before the pause idempotent, and bind each approval to a hash of the exact action.
7.  Fail closed on timeout, expire approvals and keep a kill switch that does not depend on the agent.
8.  Log who, what, tier, rule, decision, edits, timing and result, and keep what the reviewer actually saw.
9.  Sample auto-approved actions weekly for human review, and track approval rate and decision time per action type.
10.  Re-tier after every incident and every new tool: the matrix is a living document.

If you want help turning this into a concrete architecture for your agents, that is what I do in my [AI engineering work](https://balazscsorba.com/expertise/ai-engineer).

## The bigger picture

Human in the loop will not disappear as agents improve, but its shape will. The default is moving from "a person approves every step" to "a person sets the rules, reviews the exceptions and audits the rest". Auto mode becoming the default in Claude Code and the Auto-review layer in dots are two vendors reaching the same conclusion in the same season.

What stays human is accountability. A classifier can approve an action, but it cannot be responsible for it. Design every gate so that, when something goes wrong, you can answer who allowed it, on the basis of what information, and under which rule.

## Sources

1.  [LangChain docs: LangGraph interrupts](https://docs.langchain.com/oss/python/langgraph/interrupts)
2.  [LangChain docs: Human-in-the-loop (HumanInTheLoopMiddleware)](https://docs.langchain.com/oss/python/langchain/human-in-the-loop)
3.  [OpenAI Agents SDK (Python): Human in the loop](https://github.com/openai/openai-agents-python/blob/main/docs/human_in_the_loop.md)
4.  [Claude Agent SDK: Handle approvals and user input](https://code.claude.com/docs/en/agent-sdk/user-input)
5.  [Claude Agent SDK: Configure permissions](https://code.claude.com/docs/en/agent-sdk/permissions)
6.  [Claude Code docs: Hooks (PreToolUse defer, PermissionRequest)](https://code.claude.com/docs/en/hooks)
7.  [Anthropic Engineering: Claude Code auto mode](https://anthropic.com/engineering/claude-code-auto-mode)
8.  [DevOps.com: Anthropic makes Claude Code auto mode the default](https://devops.com/anthropic-makes-claude-codes-auto-mode-the-default-betting-automation-beats-manual-review/)
9.  [EU AI Act, Article 14: Human oversight](https://artificialintelligenceact.eu/article/14/)
10.  [OpenAI: How we build safety, security and privacy into dots](https://openai.com/index/how-we-build-safety-security-and-privacy-into-dots/)

## Frequently asked questions

What does human in the loop mean for AI agents?

It means a person can intervene at defined points while an agent works: approving or editing a planned action, answering a question, or stopping a run. In practice it is a gate in the agent loop. The agent pauses, shows what it intends to do, waits for a decision and continues with the outcome.

Which agent actions should require human approval?

Actions that are hard to undo, touch money, permissions or personal data, leave your organisation, or affect many records at once. Reads inside the agent scope and reversible changes in a sandbox can usually run automatically, with logging.

How do you avoid approval fatigue?

Ask less often and show more when you do. Auto-approve low-risk actions, batch related steps into one decision, show the effect and a diff instead of a tool name, block the dangerous cases by rule instead of by prompt, and audit a sample of what ran without asking.

How does interrupt and resume work in LangGraph?

A node calls interrupt with a JSON-serialisable payload. LangGraph saves the state through a checkpointer and stops. You resume on the same thread\_id with Command(resume=value), and that value becomes the return value of interrupt. The node restarts from its beginning, so side effects before the interrupt must be idempotent.

How do the OpenAI and Claude agent SDKs handle tool approval?

In the OpenAI Agents SDK a tool declares needs\_approval, the run returns the pending calls as interruptions, and you approve or reject them on a serialisable RunState and run again. In the Claude Agent SDK a canUseTool callback receives each call that no rule or mode has settled and returns allow or deny.

Does the EU AI Act require human oversight of AI agents?

Article 14 requires effective human oversight for high-risk AI systems, including awareness of automation bias and the ability to intervene or stop the system. Whether your agent is high-risk depends on its use case. Even where it is not, the same design principles are a sound baseline.

Written by Balázs Csorba

Senior fullstack & AI engineer in Styria, Austria – 10+ years of Vue, Nuxt, Node.js and PHP, now building tooling for AI agents.

[AI engineering & MCP servers →](https://balazscsorba.com/expertise/ai-engineer)[About me →](https://balazscsorba.com/about)

## More articles

-   [OpenAI dots: what always-on agents will change, and what they will not](https://balazscsorba.com/blog/openai-dots-always-on-agents-impact)
-   [Designing memory for AI agents: tiers, write rules, poisoning and GDPR](https://balazscsorba.com/blog/ai-agent-memory-design)
-   [Multi-agent systems: when they beat one agent, and when they do not](https://balazscsorba.com/blog/multi-agent-systems-when-worth-it)
-   [Building voice agents: realtime speech-to-speech or STT, LLM and TTS?](https://balazscsorba.com/blog/voice-agents-realtime-latency)

## Sounds like what you need?

Tell me about your project or role – I’d love to hear from you.

[Book a call](mailto:contact@balazscsorba.com) [Connect on LinkedIn](https://www.linkedin.com/in/balazs-csorba)
