Doherty threshold
Definition
The idea that productivity rises sharply when a system responds fast enough (usually cited as under 400 ms) that neither user nor computer waits.
The Doherty threshold is named after Walter J. Doherty of IBM, who with Arvind J. Thadhani wrote a 1982 IBM paper, The Economic Value of Rapid Response Time. Its opening line gives the idea: "When a computer and its users interact at a pace that ensures that neither has to wait on the other, productivity soars." The figure usually attached to it today, 400 milliseconds, comes from how the paper is summarized, for example on Laws of UX.
Why it matters
Before this work, a common rule of thumb was that up to two seconds of response time was acceptable, because users were assumed to be thinking about their next step while they waited. The paper argues that the facts do not bear this out, and that productivity rises faster than response time falls. It reports Thadhani's finding that a programmer completed about 180 transactions per hour at a three-second response time, and 371 at 0.3 seconds.
The broader point holds up: delays break concentration. Jakob Nielsen's response-time limits make a similar case, with 0.1 second as the limit for feeling "that the system is reacting instantaneously" and 1 second for keeping "the user's flow of thought" uninterrupted.
How to apply it
Do:
- Acknowledge every input immediately (button state, typed character, selection) even if the work behind it takes longer.
- Use optimistic UI for low-risk actions: show the result at once and reconcile with the server afterward.
- Preload and cache likely next steps so common paths feel instant.
- When work will take longer, show progress instead of a frozen screen (see AI loading states).
Don't:
- Treat 400 ms as a precise law. It is a rough target from 1980s terminal work, not a constant of human perception.
- Hide slowness behind long animations. A transition that adds 600 ms to every action makes the product feel slower, not faster.
Common mistakes
- Misquoting the source. The 400 ms figure is widely credited to Doherty and Thadhani, but the paper's text argues for subsecond response in general and reports measurements rather than stating a single threshold. Cite it carefully.
- Measuring the server, not the user. Fast API times mean little if the interface takes another second to render.
- Optimizing averages. A fast median with a slow long tail still breaks flow for many sessions.
In AI products
Language models rarely finish a useful answer in 400 ms, so AI products cannot meet the threshold for the whole response. What they can do is meet it for feedback. Acknowledge the request instantly, show a working state, and stream text as it is generated. OpenAI's latency guide says streaming cuts the waiting time "to a second or less" and that there is "a huge difference between waiting and watching progress happen."
For features that sit inside typing, such as autocomplete or inline suggestions, the original threshold applies much more directly. Suggestions that arrive after the user has moved on are noise, so these features often favor smaller, faster models over larger, slower ones.