AI error recovery
Definition
Designing for the moments an AI feature fails, so people notice the problem, understand it and can correct, retry or take over without losing work.
AI error recovery, sometimes called graceful failure, is the set of patterns that keep an AI product useful when the model gets something wrong. Classic software errors are usually crashes or invalid input. AI errors are different: the system often keeps running and confidently produces the wrong thing. Google's PAIR guidebook puts it this way: "AI is probabilistic by nature, and like all systems, will fail at some point" (PAIR).
Because failure is certain, PAIR argues "The trick isn't to avoid failure, but to find it and make it just as user-centered as the rest of your product."
Why it matters
AI errors come in kinds users can't always see. PAIR distinguishes context errors, where the system makes "incorrect assumptions about what the user wants", system limitations, where it can't provide the right answer at all, and background errors, where "neither the user nor the system register an error." Each needs a different response, and the last is the hardest because nothing on screen says anything went wrong.
How errors are handled also shapes trust. PAIR notes that these moments establish or correct mental models and calibrate user trust. A clear, recoverable error can leave people with a more accurate picture of the system than an unbroken streak of lucky answers.
How to apply it
- Do make correction quick. Microsoft's HAX guidelines include "Support efficient correction" (G9): "Make it easy to edit, refine, or recover when the AI system is wrong."
- Do degrade gracefully when the system is unsure. HAX G10, "Scope services when in doubt", suggests disambiguation or reduced service rather than a confident guess, for example offering three suggestions instead of auto-completing one.
- Do put Edit, Undo and Retry next to generated content. Apple's guidelines recommend surfacing these controls near output (Apple HIG). See undo for AI actions.
- Do explain failures in plain language with a next step. Apple advises that when something goes wrong you "describe what happened in plain language and offer a clear next step."
- Do let people take over. PAIR: "When an AI system fails, often the easiest path forward is to let the user take over." For a travel-booking agent, that means handing back a half-filled itinerary, not a blank error. Sometimes the right takeover is a person, see AI-to-human handoff.
- Don't discard the user's input or partial results when a generation fails.
- Don't show a bare "Something went wrong" for blocked requests. Apple recommends coaching people on how to get a better result next time.
Common mistakes
- Designing only for crash-style errors and ignoring wrong-but-plausible output.
- Retry buttons that produce the same failure, with no way to change the request or switch to a manual path.
- Collecting no signal. PAIR recommends feedback opportunities both on error messages and alongside correct output.
- Hiding the manual way of doing the task once AI is added, so users have nowhere to go when it fails.