Explainability
Definition
How well people can understand why an AI system produced an output. Often paired with interpretability: what that output means for their task.
Explainability is the degree to which people can understand why an AI system did what it did. In product design it covers the explanations you show: an attribution like "Because you watched X", a list of sources, a note about which inputs mattered, or a statement of what the system can't do. The goal is not to teach users machine learning but to give them enough to decide whether to rely on an output.
Explainability vs interpretability
The two words are often used interchangeably, and different research communities define them differently. A widely cited split comes from NIST's AI Risk Management Framework: "Explainability refers to a representation of the mechanisms underlying AI systems' operation, whereas interpretability refers to the meaning of AI systems' output in the context of their designed functional purposes."
In practice: explainability answers "how did it get this?", interpretability answers "what does this mean for me?". A loan model's feature weights are an explanation. "You were declined mainly because of a short credit history, and adding a co-signer would likely change the result" helps interpretation. Users usually need the second more than the first.
NIST's Four Principles of Explainable AI add a useful test for any explanation you design: it should exist, be meaningful to the intended audience, accurately reflect the system's process ("explanation accuracy"), and respect knowledge limits, meaning the system only operates under the conditions it was designed for and when it is sufficiently confident in its output.
Why it matters
Explanations shape trust. Google's PAIR guidebook frames the goal as helping users calibrate their trust, neither trusting the system everywhere nor dismissing it. An explanation that is wrong or vague can push trust in the wrong direction, which is worse than no explanation.
How to apply it
- Do prefer partial explanations. PAIR puts it plainly: "The best explanation is likely a partial one." Show the key input or data source, not the whole model.
- Do explain at the moment of action. PAIR notes that "the perfect time to show explanations is in response to a user's action", for example when a writing assistant rewrites a paragraph differently than expected.
- Do point to evidence users can check, such as sources placed next to the claim they support.
- Don't present generated step-by-step text as a faithful account of the model's process (see AI reasoning disclosure).
- Don't use first-person, human-sounding explanations ("I thought about your problem"). NN/g recommends neutral wording such as "This answer is based on the following source: [link]."
Common mistakes
- Explaining the technology instead of the outcome ("powered by a neural network" tells users nothing about this result).
- Citations that look authoritative but don't support the claim. NN/g reports that people rarely click citation links, so a fake or irrelevant source still raises confidence.
- One explanation for all audiences. A compliance reviewer and a casual user need different depth; use progressive disclosure to serve both.
- Treating explanations as a fix for an unreliable model. They help people judge outputs, they don't make outputs correct.
Sources
- NIST (2023). Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1
- Phillips, P. J. et al. (2021). Four Principles of Explainable Artificial Intelligence. NISTIR 8312
- Google PAIR People + AI Guidebook: Explainability + Trust
- Nielsen Norman Group: Explainable AI in Chat Interfaces