Tools/Security & compliance
Guardrails AI: validating what the model returns
A review of Guardrails AI: 65 hub validators, eight on-fail actions, the August 2026 shutdown of hosted inferencing, and when NeMo Guardrails fits better.
- Type
- Output validation
- Pricing
- Apache-2.0
Balázs Csorba··10 min read
- Guardrails
- Output validation
- LLM reliability
- Python

Key takeaways
- The hub ships 65 validators as individual guardrails-ai-* packages, and every validator declares one of eight failure actions, from noop to refrain.
- Only noop and exception work with a streaming response; reask, fix, fix_reask, filter and refrain need the complete output before they can act.
- Hosted remote inferencing was discontinued with a cutoff of 25 August 2026, which leaves the framework entirely local and forces a migration for anything that depended on it.
- Reask is a second full model call per attempt, so the licence is Apache-2.0 while the retry loop is the line item that grows with the failure rate.
- OpenAI lists omni-moderation-latest as free, so the case for Guardrails rests on rules, structure and audit trail rather than on moderation alone.
Guardrails AI is an open-source Python framework that checks what a model returns before an application is allowed to act on it. Validators live in a hub of 65 packages, each declaring what should happen when it fails: raise, repair, reask, filter or return nothing. The project is Apache-2.0, at version 0.11.0 published on 14 August 2026, and it is driven from Python, with JavaScript listed as supported in the README.
It occupies the layer between an LLM call and the code that consumes its output, which is where a leaked system prompt, a PII-bearing summary or a malformed JSON blob either gets caught or reaches a user. It competes with prompt-injection defence patterns for the same slot as NVIDIA NeMo Guardrails, with OpenAI's moderation endpoint as the hosted shortcut, and with the hand-written checks most teams already have. The position taken here: it is the most complete validator library in the field, and its failure-handling model is better engineered than its operating model, where the free part is the software and the expensive part is the retries.
What it actually is
The unit of work is a Guard: an ordered list of validators applied either to the messages going in, with on="messages", or to the response coming out. Validators come from the hub and install as their own PyPI packages, so a deployment only carries the checks it runs. Some are rules such as regex, length and JSON parseability, some run a small local model, and some call a second LLM to grade the first.
- Hub. 65 validators grouped by risk: brand risk, formatting, etiquette, jailbreaking, data leakage, code exploits and factuality.
- Packaging. Each validator is a separate package installed after guardrails configure, for example guardrails-ai-regex-match or guardrails-ai-detect-pii.
- Two directions. Guards run on the messages before the call and on the model output after it, and the same guard object can do both.
- Structured generation. Guard.for_pydantic drives the model against a Pydantic class and validates the parsed object, using function calling where the model supports it and prompt scaffolding where it does not.
- Licence and version. Apache-2.0, Python 3.10 to 3.13, version 0.11.0 published on 14 August 2026, with about 7,500 stars on GitHub.
- Server mode. guardrails start runs a Flask service that exposes each guard behind an OpenAI-compatible base URL, so an existing client changes one string.
How it works
Validation is synchronous by default: the raw output goes in, every validator runs in order, each failure is appended to guard.history.last.failed_validations, and the on-fail action declared on that validator decides what leaves the guard. The action is set per validator rather than per guard, so one guard can raise on PII and merely log a formatting slip. reask rebuilds a prompt containing the failed criterion and calls the model again, up to the num_reasks limit.
Because failures are recorded whether or not they stop the flow, a guard set to noop still produces an audit trail, which matters given that noop is the default. For calls that talk to a provider, the framework retries connection errors, rate limits and timeouts with exponential backoff up to a sixty-second wait, so a provider incident surfaces as latency inside the guard rather than as an immediate exception.
When a validator fails
Eight actions are available and they are the real interface of this tool, more than the validator list is. The documentation table marks which of them work against a streaming response, and that column decides more architecture choices than the feature list does.
| Action | What it does | Streaming | Where it fits |
|---|---|---|---|
| noop | Records the failure and returns the output unchanged; this is the default | Yes | Measuring how often checks fail |
| exception | Raises so the caller handles the failure | Yes | Input validation and strict pipelines |
| reask | Rebuilds the prompt with the failed criterion and calls the model again | No | Soft failures a second pass can fix |
| fix | Applies the validator's repair value, such as anonymised PII | No | PII scrubbing and formatting fixes |
| fix_reask | Fixes first, then reasks if the fixed value still fails | No | Repairs that may be incomplete |
| filter | Drops the failing field and returns the rest of a structured object | No | Structured data with optional fields |
| refrain | Returns nothing when the output is unsafe to ship | No | Content that must not reach a user |
| custom | Runs your own function over the value and the failure result | No | Policy that already lives in your code |
The opinionated part: reask is the advertised feature and the one to budget carefully, because every reask is a second complete generation billed at the same rate as the first, and it multiplies by the failure rate rather than by traffic. The documentation itself steers complex cases away from it, recommending an exception-based approach once use cases grow, because a raised exception lets one handler branch on which validator failed instead of parsing what reask silently returned.
Getting started
Installation is the core package plus whichever validators the guard needs, then guardrails configure to set up credentials. A guard is built in code, not in a config file, which keeps the failure action next to the check it belongs to.
# pip install guardrails-ai guardrails-ai-detect-pii
from guardrails import Guard, OnFailAction
from guardrails_ai.detect_pii import DetectPII
guard = Guard().use(
DetectPII(pii_entities="pii", on_fail=OnFailAction.FIX)
)
result = guard.validate(
"Hello, my name is John Doe and my email is john.doe@example.com"
)
print(result.validation_passed) # True once the scrub lands
print(result.validated_output) # <PERSON> and <EMAIL_ADDRESS>For callers that are not Python, guardrails start serves each guard at a base URL shaped like localhost:8000/guards/<name>/openai/v1/, so an existing OpenAI client points at it without a new SDK. The README also lists JavaScript as supported, with the Python package remaining the reference implementation.
Cost
The software is free in the strict sense: Apache-2.0 covers the framework and every hub validator, and nothing metered sits between install and validation. The bill is elsewhere, in what the validators call on the way through.
- Licence. Apache-2.0 for the framework and the validators; there is no paid tier of the library itself.
- LLM-backed checks. Validators such as llm_critic, provenance_llm and qa_relevance_llm_eval add one model request per check per response.
- Reasks. Each reask is a full regeneration, and num_reasks sets how many times one response may be regenerated before the guard gives up.
- Local ML. Validators like detect_pii run Presidio in-process, so their price is memory and latency rather than tokens.
- Server mode. The optional Flask service is infrastructure the deployment runs, scales and secures itself.
The comparison that matters is against hosted moderation. OpenAI lists omni-moderation-latest as free on its pricing page, and NeMo Guardrails is Apache-2.0 as well, so neither competitor charges a licence for the same job. What Guardrails sells is coverage: a moderation endpoint answers one question about content policy, while the hub can also enforce schema, length, competitor mentions, prompt leakage and provenance in the same pass.
Where it shingles
The weaknesses are operational, and they show up after the first week. The hub mixes a regex, a BERT model and an LLM judge behind one abstraction, so latency differs by orders of magnitude between validators and there is no per-validator performance budget in the docs. The hub's language filter lists English only. Streaming support stops at noop and exception, which rules the tool out for token-by-token delivery unless the whole response is buffered. And the hosted inferencing shutdown in August 2026 was a breaking change to a service some deployments had built on.
| Attribute | Guardrails AI | NVIDIA NeMo Guardrails | OpenAI Moderation API |
|---|---|---|---|
| Shape | Python library plus 65 hub validators | YAML rules and Colang dialog flows | One hosted endpoint |
| Licence | Apache-2.0, entirely local after August 2026 | Apache-2.0, optional anonymous telemetry | Closed, served by OpenAI |
| Configuration | Python code, on_fail declared per validator | Files in a rails directory | A single moderation request |
| Cost per request | Free software; LLM validators and reasks bill separately | Free software; rails may call an LLM | omni-moderation-latest listed as free |
The third alternative is the one most teams actually ship: hand-written checks around the call, an if-statement for the JSON parse and a regex for anything resembling an account number. That is cheaper than all three rows above and it fails silently by construction, which is precisely the failure mode this category exists to remove. The honest comparison is not Guardrails against NeMo, it is Guardrails against whatever the team writes in an afternoon and then forgets to extend.
Verdict
Adopt it where a bad response costs money, data or credibility, and where the record of what failed matters as much as the failure itself. It is the most complete validator catalogue available under a free licence, and the per-validator failure action is a better design than a global policy switch. Its costs are the ones every in-process checker has: you run it, you tune it, and you pay for whatever calls the validators make.
- Use it when a response leaves the service: PII in a support summary, SQL a user executes, structured output another service parses.
- Use it where an audit trail is required, since every failure lands in guard.history, which is more than a hand-written check usually records.
- Prefer NeMo Guardrails when the requirement is conversational policy, such as when to decline or change topic, because that is what Colang rails are built for.
- Prefer the free moderation endpoint when abusive content is the only risk and adding a dependency is not worth it.
- Do not expect streaming fixes: only noop and exception act on a partial response, so buffer the stream or drop the tools that repair.
The documentation's own advice once a use case outgrows the simple path reads, as usecases get more complex, we recommend switching to an exception-based approach. A framework that tells you to stop using its headline feature when things get hard is being honest about where that feature belongs.
Sources
- Guardrails AI on GitHub: README, news and FAQ
- guardrails-ai 0.11.0 on PyPI
- Guardrails Hub: 65 validators
- Guardrails documentation: error remediation and on-fail actions
- Guardrails documentation: use on-fail actions
- Migration issue 1560: moving off hosted remote inferencing
- NVIDIA NeMo Guardrails on GitHub
- NVIDIA NeMo Guardrails documentation
- OpenAI API pricing, including moderation
Frequently asked questions
Is Guardrails AI free?
The framework and the validators are Apache-2.0 and install from PyPI without a licence fee. The cost appears in what the validators invoke: an LLM-backed check adds a model request per response, and reask regenerates the answer in full for each attempt.
What happens when a validator fails?
The failure is written to guard.history and then the validator's on-fail action runs: noop logs and passes the value through, exception raises, fix applies a repair such as anonymised PII, reask calls the model again, filter drops the failing field, refrain returns nothing, and a custom function receives the value and the failure result.
Does it work with streaming output?
Only noop and exception are documented as streaming-compatible. reask, fix, fix_reask, filter and refrain all require the complete output, so a token-by-token stream has to be buffered before those actions can run.
Guardrails AI or NVIDIA NeMo Guardrails?
Guardrails is a Python library of validators wrapped around a call you already make; NeMo is a runtime that owns the conversation flow through YAML rules and Colang dialogs. The first fits into existing code, the second replaces part of it.