AI guardrails
Definition
Rules, filters and permission limits around an AI system that keep its inputs, outputs and actions within safe, intended bounds.
Guardrails are the controls that sit around a model rather than inside it: system instructions that scope what it should do, filters that check inputs and outputs, and permission limits on what tools or data it can touch. They exist because a model alone can't be relied on to stay on topic, refuse harmful requests or resist manipulation every time. For designers, guardrails are also a user experience: how a blocked request, a refusal or a confirmation step feels decides whether people understand the boundary or just feel stonewalled.
Why it matters
Generative systems face risks ordinary software doesn't. NIST's generative AI profile lists risks such as confabulation and dangerous, violent or hateful content (NIST AI 600-1). OWASP ranks prompt injection first in its 2025 list for LLM applications, describing it as when "user prompts alter the LLM's behavior or output in unintended ways." Indirect prompt injection, where the model reads instructions hidden in "external sources, such as websites or files", is especially relevant to agents that browse or read email (OWASP LLM01).
The more an AI product can act, the more its guardrails matter. A chat assistant that says something wrong is a problem. An agent that sends money or deletes files because a web page told it to is a bigger one.
How to apply it
OWASP's mitigations for prompt injection map well onto product decisions:
- Do constrain behavior to the task. A support assistant for a bank should stay on banking.
- Do enforce least privilege. A meeting-notes tool needs read access to the calendar, not send rights on email.
- Do require human approval for high-risk actions. Apple's guidelines advise avoiding automating destructive actions and actions that are hard to undo, "like making a purchase on a person's behalf" (Apple HIG). See human-in-the-loop.
- Do mark external content as untrusted so the model treats it as data, not instructions.
- Do explain refusals helpfully. Apple recommends coaching people toward a better request when output is blocked.
- Do red-team with vague, out-of-scope, sensitive and adversarial prompts before launch, as Apple and OWASP both recommend.
- Don't rely on the system prompt as your only guardrail. Instructions can be overridden. Hard limits belong in code and permissions.
- Don't over-block. Refusing ordinary requests teaches users that the feature is unreliable.
Common mistakes
- Treating guardrails as a launch checklist item rather than something you monitor and update. Apple notes that some improvements, like updating a list of blocked words, can ship independently of app releases.
- Vague refusal messages that don't say what was blocked or what the user can do instead.
- Giving an agent broad API keys "for convenience", which turns any prompt injection into a real-world action.
- Filtering outputs but not inputs, or the reverse.