Human-in-the-loop
Definition
A design where a person reviews, approves, corrects or can stop an AI system's work at key points, instead of the system acting fully on its own.
Human-in-the-loop (HITL) describes an AI system that keeps a person involved at the points where judgment matters: approving an action before it runs, reviewing a draft before it is sent, correcting a classification, or stopping a process that is going wrong. In machine learning the phrase can also mean people labeling or correcting training data. In product design it usually means checkpoints in the user's workflow.
Why it matters
AI systems make mistakes, and some mistakes are expensive or hard to reverse. A human checkpoint is the cheapest place to catch them. It is also increasingly a legal expectation: Article 14 of the EU AI Act requires high-risk AI systems to be designed so that people overseeing them can, among other things, "disregard, override or reverse the output" and "interrupt the system through a 'stop' button or a similar procedure."
People also often want control. Google's PAIR guidebook lists situations where people prefer to stay in charge, including when they enjoy the task, feel personally responsible for the outcome, or when the stakes are high.
How to apply it
- Do put the checkpoint before irreversible actions. Apple's generative AI guidelines advise: "Generally, ask for confirmation before performing a significant action on someone's behalf." An agent that books travel should show the itinerary and price before paying.
- Do make the review step meaningful: show what will happen, what the AI assumed, and what the user can change, not a bare "Continue?" button.
- Do scale oversight to risk. PAIR's Explainability + Trust chapter suggests you "Progressively increase automation under user guidance" and automate more "when trust is high, or risk of error is low."
- Do give a manual fallback. PAIR calls the manual method "a safe and useful fallback."
- Don't add approvals for trivial, reversible steps. Too many prompts train people to click through, which defeats the point.
- Don't hide the stop control. Long-running agents need a visible way to pause or cancel.
Common mistakes
- Rubber-stamp review. If the human only sees the final output, not the inputs and assumptions, approval becomes a formality. The AI Act names this risk directly: overseers should remain aware of the tendency of "automatically relying or over-relying" on AI output, which it calls automation bias (see automation bias).
- Checkpoint in the wrong place. Reviewing an email after it is sent is not oversight. Put the review where a change is still cheap.
- No feedback path. When the human corrects the AI, nothing records it, so the same error repeats (see AI feedback loops).
- All or nothing. Offering only "fully manual" or "fully automatic" ignores the middle ground of suggest, draft, and act-with-confirmation.