AI feature spec (PRD) template
A normal product spec assumes the feature does what it was built to do. An AI feature won't, some of the time, so its spec needs sections a normal one doesn't: why AI rather than rules, which mistakes matter most, what happens when the model fails or is unavailable, and how the team will know the output is good enough to ship.
This template has those sections. The prompts under each heading come from Google PAIR's People + AI Guidebook, Microsoft's Guidelines for Human-AI Interaction and Apple's guidelines for generative AI, and are linked where they apply. Fill it in on this page (your answers stay in this browser), copy it as Markdown into your own docs, or download the PDF to print.
Download the PDF
AI feature spec (PRD) template as a print-ready PDF (A4, 157 KB). Everything in it is also free to read on this page.
How to use it
Write the user problem and the failure modes before anything about the model. If the failure-modes table is hard to fill in, the team isn't ready to build yet.
Keep it short. A blank you can't answer is more useful than a paragraph that hides the gap; mark it as an open question.
Answers are saved in this browser only. Nothing you type here is sent to UX Pickle.
1. Overview
What people will see it called in the product.
Product, design, engineering, and who owns model quality.
Draft, in review or approved; date of last change.
Who it's for, what it does for them, and what the AI part is.
2. User problem
The specific users or segment, and the situation they're in.
What they are trying to do and what gets in the way.
Current workaround, tool or manual process, and what it costs them.
Research, support tickets, analytics. Link them.
3. Where AI helps
PAIR: find the intersection of user needs and AI strengths, and decide whether to automate or augment. PAIR: User Needs + Defining Success
What can AI do here that a rule-based or simpler design can't? What would the non-AI version look like?
Does the AI do the task, or help the person do it? Why is that the right choice for this task?
Inputs, outputs, and where in the flow it appears.
Explicit limits. These become the “what it can't do” copy (Microsoft G1).
Entry points, defaults, and how to ignore or turn it off (Microsoft G7, G8, G17).
4. Data, model and privacy
Apple HIG: Generative AI (Privacy)
User content, account data, usage data, third-party data. Mark anything personal or sensitive.
On device, own servers, or a third-party model provider. Which provider and region.
Is user content used to train or improve models? How long are prompts and outputs kept?
What people are told, what they agree to, and how they opt out.
Where the product says AI is involved; whether EU AI Act Article 50 applies (chatbot notice, machine-readable marking, deep fakes).
5. Failure modes
List the ways it will be wrong, including errors that are “working as designed” but still wrong for the person. Weigh false positives against false negatives. PAIR: Errors + Graceful Failure
Failure modes
Start with: a wrong but confident answer (hallucination), no answer, a refusal, a slow or timed-out response, the model unavailable, harmful or biased output, and misuse.
| Failure | Example | Impact on the person | How we detect it | What the UI does |
|---|---|---|---|---|
False positive or false negative? So do we tune for precision or recall?
6. Human fallback
Apple HIG: Generative AI (Best practices)
How people complete the task when the AI is wrong, unavailable, or switched off.
How people edit, retry, revert or dismiss the output (Microsoft G9).
When and how a human takes over. Who reviews; how fast.
Does the feature make decisions with legal or similarly significant effects? If so, GDPR Article 22 safeguards: human review, a way to contest.
7. Success metrics
Metrics
Types: user outcome (task success, time saved), output quality (accuracy, groundedness, rated helpfulness), guardrail (harmful output rate, complaints, opt-outs), adoption and retention, and cost or latency.
| Metric | Type | Baseline | Target | Launch threshold |
|---|---|---|---|---|
Metrics that could rise while the experience gets worse, such as engagement driven by wrong answers that need follow-ups.
8. Evaluation plan
Test across a diverse set of people and inputs, including vague, out-of-scope, sensitive and adversarial requests. Apple HIG: Generative AI (Outputs)
Where test inputs come from, how many, how representative, and which adversarial cases are included.
Automated checks, human rating rubric, who rates, and agreement between raters.
Usability and trust testing before launch, including a wrong-answer scenario. Link the research plan.
Feedback signals, error reports, sampling and review cadence, alert thresholds.
What runs before a model, prompt or provider change ships; how users are told about noticeable changes (Microsoft G18).
9. Launch and open questions
Staged rollout, who gets it first, and the switch to turn it off.
Product, legal, reputational and accessibility risks.
Anything the team can't answer yet, with an owner and a date.
Sources
- Google PAIR: People + AI Guidebook
- Microsoft HAX Toolkit: Guidelines for Human-AI Interaction
- Apple Human Interface Guidelines: Generative AI
- EU AI Act, Article 50 (EUR-Lex)
- GDPR, Article 22 (EUR-Lex)
- PAIR: User Needs + Defining Success
- PAIR: Errors + Graceful Failure
Checked against these sources on 3 October 2026. Spotted something out of date? Email hi[at]uxpickle.com.