AI UX patterns

Confidence indicators

Definition

UI cues that show how certain an AI system is about an output, such as labels, ranked alternatives or ranges, so people know how much to rely on it.

Confidence indicators communicate uncertainty in an AI output. They range from an explicit label ("High confidence") or a percentage, to softer signals such as showing several alternatives, hedged wording ("we think you'll like"), or simply not showing a result when the model is unsure. Displaying uncertainty well is one of the main tools for helping users rely on AI the right amount.

Why it matters

A model that sounds equally sure about everything invites over-reliance on its weak answers and under-use of its strong ones. Microsoft's Guidelines for Human-AI Interaction ask designers to "Make clear how well the system can do what it can do" (G2), and confidence display is the per-result version of that. But confidence UI can also mislead. Google's PAIR guidebook warns that "A misleadingly high confidence, for example, may cause users to blindly accept a result," and that showing more granular confidence "can be confusing if the impact isn't clear."

Common formats

PAIR lists four main visualizations:

  • Categorical: buckets like High / Medium / Low. Your team sets the cutoffs, and each category should map to a clear user action.
  • N-best alternatives: "This photo might be of New York, Tokyo, or Los Angeles." Useful when confidence is low, because it prompts users to use their own judgment.
  • Numeric: a percentage. Risky, because it assumes users understand probability. PAIR notes that novice users may not know whether 80% is high or low in context.
  • Data visualizations: error bars or shaded ranges, best suited to expert users in domains that already use them.

How to apply it

  • Do confirm that confidence scores track real accuracy first. Apple's machine learning guidelines say that if you're not sure how confidence values correlate with result quality, it's not a good idea to show them.
  • Do translate scores into something actionable. Apple suggests categories like "high chance" and "low chance" for price predictions, or advice like "This is a good time to buy."
  • Do change behavior at thresholds, not just labels. Apple's example: Photos shows matches directly when confidence is high, and asks people to confirm when it is lower.
  • Do set a floor for proactive features. A meeting-notes summariser that isn't sure who owns an action item can leave the owner blank and ask, rather than guess.
  • Don't show "97% match" when users can't act on the difference between 97 and 92.
  • Don't rely on token-level probabilities from a language model as a user-facing truth score without validating them on your own tasks.

Common mistakes

  • Adding confidence badges without testing whether they change decisions. PAIR advises setting aside time to test whether showing confidence helps at all.
  • Using the same visual weight for low-confidence and high-confidence results.
  • Showing precise numbers to a general audience where a ranked list or a hedge would do.
  • Forgetting that a confident tone in generated text is itself a confidence signal, often an unearned one.

Sources

  1. Google PAIR People + AI Guidebook: Explainability + Trust
  2. Apple Human Interface Guidelines: Machine learning
  3. Amershi, S. et al. (2019). Guidelines for Human-AI Interaction. CHI 2019

Browse the full AI UX glossary