AI error, hallucination and failure-state checklist

Checklist · 27 items · Updated

Every AI feature is wrong some of the time. Language models add a specific kind of error: a hallucination, an answer that reads as confident and plausible but is made up. Apple's guidelines describe it as content that “seems plausible but is made up”, and NN/g points out that the model has no way of knowing whether its output is true.

So the design question isn't whether the feature will fail, but whether people notice when it does and can get to what they needed anyway. This checklist covers the errors to plan for, how to communicate uncertainty, the failure screens most AI features need, and how people recover.

Google PAIR's Errors + Graceful Failure chapter is the backbone; Apple's Human Interface Guidelines and NN/g's research fill in the generative-AI specifics.

Download the PDF

AI error, hallucination and failure-state checklist as a print-ready PDF (A4, 136 KB). Everything in it is also free to read on this page.

How to use it

List the feature's failure modes first (the first section), then check each failure state below against the real product. Trigger them on purpose: turn off the network, send an empty or very long input, ask something out of scope, and ask something you know the answer to but the model probably doesn't.

0 of 27 done

Ticks are saved in this browser only. Nothing you tick here is sent to UX Pickle.

Know your errors

  • PAIR notes that people judge errors by their expectations: something working as designed can still feel like a failure. It calls these context errors. PAIR: Errors + Graceful Failure

  • PAIR groups AI errors by source because the fix differs: better data, better input handling, or better fallbacks. PAIR: Errors + Graceful Failure

  • PAIR asks teams to account for the situation's stakes. A wrong song is harmless; a wrong dosage, deadline or amount is not. PAIR: Errors + Graceful Failure

  • Apple recommends avoiding requests for factual information unless the model has access to verified, up-to-date information, and avoiding AI-generated content where a hallucination could misinform and harm. Apple HIG: Generative AI

  • Apple recommends testing exactly these cases and challenging your own policies and expected use cases. Apple HIG: Generative AI

Communicating uncertainty

  • Apple asks apps to clearly communicate that AI-generated content may contain errors. NN/g recommends placing disclaimers near the input and pairing them with an action. Apple HIG: Generative AI

  • NN/g lists options: uncertain wording, the factors behind a prediction, confidence ratings, multiple responses, and visible sources, and warns that generic disclaimers are easily ignored. NN/g: AI Hallucinations: What Designers Need to Know

  • Sources let people verify a claim instead of trusting tone. Make checking cheap: link to the passage, not the site. NN/g: AI Chatbots Discourage Error Checking

  • NN/g found that polished formatting and well-cited appearance lead people to over-trust output and skip checking it. NN/g: Explainable AI in Chat Interfaces

  • Microsoft's G10: scope services when in doubt; disambiguate or degrade gracefully when uncertain about the user's goal. Microsoft HAX: Guidelines for Human-AI Interaction

  • NN/g cautions that such walkthroughs are often rationalisations generated after the fact, not faithful accounts of the model's process. NN/g: Explainable AI in Chat Interfaces

Failure states to design

Each of these needs its own message and next step. A single “Something went wrong” covers none of them well.

  • PAIR calls these failstates and recommends giving people a path forward, including returning control to them. PAIR: Errors + Graceful Failure

  • Apple recommends coaching people to be more successful next time, and offering example requests that lead to better results. Apple HIG: Generative AI

  • Apple notes generative models take longer and suggests designing a loading experience or generating in the background. Apple HIG: Generative AI

  • Losing a long prompt or a half-written draft turns a temporary fault into lost work.

  • An incomplete answer that looks complete is a silent error.

  • Apple notes models may be unavailable because of device compatibility, network access or battery, and recommends a non-AI fallback where possible. Apple HIG: Generative AI

  • A limit people couldn't see coming reads as a broken product.

  • Apple notes it may not be possible to prevent every harmful outcome and recommends policies and evaluation to minimise them. Apple HIG: Generative AI

Recovery and correction

  • Apple and Microsoft (G9, efficient correction) both ask for easy ways to refine or revert results. Apple HIG: Generative AI

  • PAIR recommends returning control to people when the AI fails and always providing a manual alternative. PAIR: Errors + Graceful Failure

  • Apple recommends acknowledging when corrections take effect, which helps people build an accurate mental model. Apple HIG: Generative AI

  • An AI that can't help and can't hand over leaves people stuck at the moment they most need help.

Learning from errors

  • PAIR treats feedback on errors as a way for people to help improve the system. PAIR: Errors + Graceful Failure

  • PAIR recommends monitoring multiple channels to find new errors as people use the product in ways you didn't test. PAIR: Errors + Graceful Failure

  • Generative output can change with small input or model changes; a fixed test set is how you notice.

  • Microsoft's G18: notify users about changes; G14: update and adapt cautiously. Microsoft HAX: Guidelines for Human-AI Interaction

Sources

Checked against these sources on 3 October 2026. Spotted something out of date? Email hi[at]uxpickle.com.