AI UX patterns

AI loading states

Definition

What an AI product shows while a model is working, from spinners to step-by-step status, so waits of seconds or minutes feel clear rather than broken.

AI loading states cover the time between a user's request and a usable result. Model latency varies far more than a typical page load: a short reply may arrive in under a second, a long report or an agent run can take minutes. The loading state has to tell people that the system is working, roughly how long it will take, and ideally what it is doing.

Why it matters

Classic response-time research still applies. Jakob Nielsen's three limits put 1 second as "the limit for the user's flow of thought to stay uninterrupted" and 10 seconds as "the limit for keeping the user's attention focused on the dialogue." Many generative tasks blow past both.

Apple's generative AI guidelines note that "Generative models typically take longer to produce a result, so design a loading experience or generate in the background while a person uses another part of the app." They also recommend specific status text: instead of "Processing…", say something like "Summarizing key themes from your notes," because "Specific feedback reduces uncertainty and makes waiting feel purposeful."

How to apply it

Match the indicator to the expected wait:

  • Under about 1 second: no indicator. NN/g's skeleton screens guidance says indicators would be counterproductive here.
  • A few seconds: a spinner or animated placeholder, or a skeleton for full-page content. NN/g suggests looped animation for waits of roughly 2-10 seconds.
  • Text answers: stream the response. OpenAI's latency guide calls streaming "the single most effective approach."
  • Long or multi-step work: show real steps ("Searching 3 sources," "Reading your calendar," "Drafting summary") and, where possible, progress. OpenAI's guide advises: "If you're taking multiple steps or using tools, surface this to the user."
  • Over 10 seconds: give an estimate or percent-done indicator, and let people leave and come back. NN/g reports that in one study, people who saw a moving progress bar "were willing to wait on average 3 times longer" than those who saw no indicator.

Do: keep the user's request visible while waiting, offer Cancel, and notify when a background task finishes.

Don't: freeze the input with no explanation, or show a generic "Thinking…" for a minute.

Common mistakes

  • Fake progress. Status messages that cycle on a timer, unrelated to what the system is doing, are misleading once users notice.
  • One indicator for every wait. A full-screen loader for a two-second suggestion is as wrong as a tiny spinner for a three-minute agent run.
  • No timeout state. If the model stalls, say so and offer a retry instead of spinning forever.
  • Blocking the whole app. Long generations should not stop people from reading or working elsewhere.
  • Ignoring the first-token gap. Even with streaming, the pause before the first word needs a visible working state.

Sources

  1. Apple Human Interface Guidelines: Generative AI
  2. OpenAI API docs: Latency optimization
  3. Nielsen, J. (1993). Response Times: The 3 Important Limits. Nielsen Norman Group
  4. Nielsen Norman Group: Progress Indicators Make a Slow System Less Insufferable
  5. Nielsen Norman Group: Skeleton Screens 101

Browse the full AI UX glossary