Guardrails
Letter GDeliberate limits on what a system can say, do, or cause. In AI they constrain the model; in experimentation they flag when a gain degrades an outcome you cannot afford to lose.
Updated:
Working definition
A guardrail is a deliberate limit on what a system can say, do, or cause. It is defined before optimization: it protects outcomes that cannot be sacrificed even if another metric or capability improves.
In AI products, guardrails are the rules, filters, and permissions that constrain a model before, during, and after generation. Asking a model to “be careful” is not a guardrail. A guardrail is applied, observed, and, when possible, measured.
Two uses of the term
The same name shows up in two different practices:
- In AI: limits on content, actions, and tools. Which topics are out of scope, which data must not leave, which operations the model cannot trigger on its own.
- In experimentation: guardrail metrics. If a variant improves the primary goal but degrades an outcome we cannot afford to lose — errors, trust, conversion, response time — the experiment is not treated as a success.
The shared idea is the same: do not judge an improvement by a single variable.
The article UX guardrails develops the product-limit sense: permissions, confirmation, and recovery. How to make design decisions uses the second sense when it distinguishes goal metrics, guardrails, and data-quality checks.
Why it matters in UX
An AI interface is not a generic chat. It is a product that speaks, recommends, and sometimes acts. Without explicit limits, the model can present an invention as a fact, leave the product’s scope, trigger an action the person did not ask for, or expose data that should not leave.
It can also fail the other way: refuse opaquely, without anyone understanding what happened or how to continue.
UX work does not end with the assistant’s tone. It includes designing the boundaries: what the system may do, how a refusal is communicated, what evidence it shows, and what remains on a person’s side.
Three moments in an AI product
- Before generation: filter inputs, detect sensitive data, bound topics, and limit which documents are retrieved.
- During generation: system policies, allowed tools, and restrictions on which actions the model can start.
- After generation: classify the output, redact sensitive information, and require citations or human review when the risk justifies it.
A single filter at the end is not enough. Risk appears in the input, in the tools, and in what is shown as a response.
Relationship with evals
Evals do not replace guardrails: they check whether the guardrails hold. A limit nobody tests is a policy. A limit with failure cases, thresholds, and periodic review is a control.
Without evals, teams usually discover the holes once people are already on the other side.
Anti-patterns
- Prompt as the only layer: asking the model not to fail, without filters, permissions, or tests.
- Over-blocking: constant refusals destroy usefulness and push people toward workarounds.
- Invisible refusal: the system does not do something and does not explain what happened or how to continue.
- Optimizing a single metric: improving speed, conversion, or “perceived quality” without watching errors, trust, or compliance.
- Treating brand voice as safety: the right tone does not prevent the wrong action.
Quick check
- What must this system not say, do, or cause, even if it “works”?
- Does that limit live in the input, the tools, the output, or only in a paragraph of the prompt?
- Can the person understand a refusal and continue?
- Which metric would warn us that an improvement is breaking something we cannot lose?
- Are there evals for the failures that actually matter?
Short definition
A guardrail is a limit set in advance to protect what cannot be sacrificed. In AI, it constrains what a model can say or do. In experimentation, it flags when an improvement degrades a critical outcome.