AI UX research and usability test plan template

Template · 8 sections, 27 fields · Updated

Testing an AI feature is not only about whether people can use it. It is also about whether they trust it the right amount: relying on it where it is good, checking it where it is weak, and noticing when it is wrong. Microsoft's literature review on overreliance defines the problem simply: overreliance “occurs when users start accepting incorrect AI outputs”. Google PAIR calls the goal calibrated trust.

This plan adds what a standard usability test leaves out: a session setup that controls what the model says, a scenario where the AI gives a plausible wrong answer, and questions that compare how much people trust the output with how good it actually is.

Fill it in on this page (answers stay in this browser), copy it as Markdown, or download the PDF.

Download the PDF

AI UX research and usability test plan template as a print-ready PDF (A4, 139 KB). Everything in it is also free to read on this page.

How to use it

Control the outputs. If participants see different live model answers, you can't compare sessions. Use a prototype with scripted responses (a Wizard of Oz setup), a fixed set of cached outputs, or a pinned model version with logged outputs.

Plan the wrong-answer scenario with care. Tell participants at the end that one answer was wrong on purpose, and why.

0 of 27 filled in

Answers are saved in this browser only. Nothing you type here is sent to UX Pickle.

1. Study overview

What it is, and the build: concept, prototype with scripted outputs, or live model (with version).

The decisions this study will inform.

For example: Do people understand what the feature can and can't do? Do they notice a wrong answer? Can they correct it? Do they know where its answers come from?

Moderated or unmoderated, remote or in person, session length.

2. Participants

Target users, plus a mix of experience with AI tools (none, occasional, daily), because trust varies with it.

How many people in each segment.

3. Setup

Scripted, cached or live; which answers are correct and which one is wrong on purpose.

Realistic content for the feature to work on, with nothing personal. How it is reset between sessions.

What the moderator does if a live model produces something unexpected or harmful.

4. Before first use: expectations

Ask before people try the feature, so their first impression isn't shaped by the output (Microsoft G1, G2). Microsoft HAX: Guidelines for Human-AI Interaction

Compare with what it actually does.

Record their estimate; you'll compare it with the true rate later.

Checks the mental model of data use.

5. Tasks

Tasks

Cover: a first use with no instructions; the core task; checking a sourced answer; correcting or refining an output; dismissing or switching off a suggestion; recovering from a failure state (timeout, refusal or no answer).

TaskScenario given to the participantSuccess criteriaWhat to observe

6. Wrong-answer scenario

One task includes a plausible but wrong output. Did the participant notice, check it, and recover? Overreliance is accepting incorrect AI output. Microsoft Research: Overreliance on AI, Literature Review (Passi, Vorvoreanu, 2022)

Exactly what the AI says, and why it is wrong. Plausible, not absurd: a wrong date, figure or step.

What would happen to the person if they acted on it.

Noticed (yes/no, when), checked a source (yes/no, how), corrected (yes/no, how), said anything about trust.

Reveal the planted error. Ask: “What would have helped you notice?” and “Does this change how you'd use it?”

7. Trust calibration questions

Ask after specific answers, not only at the end. Compare stated trust with whether the answer was actually right. PAIR: Explainability + Trust

Use a 1–7 scale and ask why.

Shows whether people know where the feature is weak.

Tests their mental model of the system (Microsoft G11).

Checks whether uncertainty signals and citations are understood as intended.

Compare with the pre-use estimate and the true rate in the session.

8. Metrics and analysis

Metrics

Task success; time on task; wrong-answer detection rate; overreliance (accepted a wrong output); under-reliance (rejected or redid a correct output); corrections attempted and completed; post-task ease rating.

MetricHow it is measuredResult

How findings will be grouped (understanding, reliance, correction, failure recovery) and who reviews them.

Format, audience and date.

Sources

Checked against these sources on 3 October 2026. Spotted something out of date? Email hi[at]uxpickle.com.