Streaming responses
Definition
Showing an AI model's output as it is generated, token by token, instead of waiting for the full answer before displaying anything.
Streaming means rendering a language model's answer while it is still being written. Large models can take many seconds to finish a long response, but the first words are usually ready much sooner. Instead of a spinner followed by a wall of text, the user sees the answer grow on screen. Most model APIs support it: both OpenAI and Anthropic deliver streamed output as server-sent events when you set a stream flag on the request.
Why it matters
Waiting with no feedback feels longer than watching work happen. OpenAI's latency guide calls streaming "the single most effective approach" to making users wait less, because it cuts the waiting time "to a second or less." The same guide points out that streaming is not only a perception trick: the user can start reading straight away, so they finish reading sooner than if the whole answer appeared at once.
Streaming also lets people judge an answer early. If the first sentence shows the assistant misunderstood the question, the user can stop it and rephrase instead of waiting for the full, wrong response.
How to apply it
Do:
- Stream any response that takes more than a moment to generate, such as a chat assistant's reply or a meeting-notes summary.
- Give users a visible Stop button while text is streaming, and keep the partial output when they press it.
- Keep the reading position stable. NN/g's chatbot guidelines advise to "keep the user's scroll position at the top of the new message rather than jumping to the bottom," because autoscrolling a streaming answer pulls people away from where it starts.
- Show a brief working state before the first token arrives, so the gap between send and first word is not blank (see AI loading states).
- Plan for the stream to fail partway. Anthropic's docs note the API "may occasionally send errors in the event stream," such as an overloaded error. Show what arrived, mark it as incomplete, and offer a retry.
Don't:
- Reflow the layout on every token. Jumping line breaks and tables that resize as they fill are hard to read.
- Enable actions like Copy, Insert into document or Send while the response is still incomplete, or make clear that they act on partial text.
- Stream output you must check first without a plan. OpenAI warns that streaming "makes it more difficult to moderate the content of the completions, as partial completions may be more difficult to evaluate." Its latency guide suggests processing output in chunks on your back end before forwarding them.
Common mistakes
- Treating streaming as a substitute for speed. A fast first word does not help if the useful part arrives 40 seconds later. Put the answer first in the response format.
- Streaming structured output raw. Half-written JSON, code fences or Markdown tables render badly. Buffer structured blocks until they are complete, or render them with a placeholder.
- Losing the partial answer on error. If the connection drops, users should not lose text they were already reading.
- No way to stop. A long, wrong answer that cannot be interrupted wastes time and tokens.
- Animating fake typing. Adding artificial character-by-character delays to text that is already complete slows people down for no benefit.