Turbo AI PM
Turbo AI PM

How to Write an AI PRD

Traditional PRDs index on user stories and pixel-perfect UI. AI PRDs need to index on constraints, edge cases, context, and safety bounds. This guide covers every section a traditional PRD skips — and why each one matters more than the feature description.

This is a framework, not a template to copy verbatim. Every AI product has different context requirements, safety needs, and team dynamics. Adapt aggressively — the goal is a document your engineering team can actually build from.
01

How an AI PRD differs from a traditional one

A traditional PRD describes what the system does in deterministic terms. An AI PRD has to describe what the system is allowed to do, what it must never do, what information it needs, and what happens when it's wrong — none of which fits cleanly into a user story.

Traditional PRD focuses on
AI PRD must also cover
User stories"As a user I want to…"
Behavioral constraintsWhat the model must and must never do, regardless of what the user asks
UI specificationsWireframes, flows, states
Context specificationExactly what data the model needs to do its job — and what it must not have access to
Acceptance criteriaBinary pass/fail for deterministic behavior
e.g. "If user clicks X, Y happens"
Eval criteria & directional boundsThresholds, slices, baselines, and the definition of "good enough" — written as bounds, not exact strings. Instead of "model says 'Hello John'", write "model greets user by name and maintains a professional tone." See Go/No-Go Rubric.
Error statesNetwork timeout, invalid input
Failure modesHallucination, over-refusal, bias, confidence-score behavior, graceful degradation design
Launch criteriaQA sign-off, load test passing
Safety sign-offExplicit unforgivables defined, red-team run, human-review path in place
Post-launch monitoringUptime, error rate
Drift monitoring + flywheel designHow production signal feeds back into the next eval set (see Data Flywheel)
The most important shift In a traditional PRD, the PM describes what the product does and engineering figures out how. In an AI PRD, the PM also has to describe the operating envelope — the bounded space inside which the model is allowed to act. Without that, engineering will make those calls implicitly, and the PM will find out when the CEO sends a Slack message at 11pm.
Writing AI acceptance criteria: Never use exact string matches. Write directional bounds — behavior the output must satisfy, not words it must contain. "Model greets user by name and maintains a professional tone, without referencing account details it was not given" is AI acceptance criteria. "Model says 'Hi [name], how can I help?'" is not — it will fail on every response that isn't that exact string.
02

Defining the context

The single most underspecified part of most AI PRDs. "The model needs access to user data" is not a context specification. A proper context spec tells engineering exactly what information flows into the prompt, why, and what doesn't — and it forces the PM to think through privacy, latency, and cost implications before code is written.

User context

Who the user is

Identity, permissions, role, subscription tier, language/locale, preferences, past behavior relevant to this task.

e.g. "The model must know the user's current subscription tier and whether they have admin permissions. It must not have access to payment card details."
Session context

What's happening right now

Current task state, recent conversation history, page/product context, what the user just did.

e.g. "Include the last 3 messages in conversation history. Truncate beyond that with a summary. The model must know which product page the user is currently viewing."
Business context

The company's data and rules

Relevant policies, product catalog, knowledge base, pricing rules, SLAs, or any internal data that changes what the correct answer is.

e.g. "Retrieve the top 3 relevant policy documents via RAG. The model must use the current pricing as of today's date — never cached pricing."
For every piece of context: define it, explain why it's needed, specify where it comes from, set a staleness limit, and explicitly list what is NOT included and why.
Context elementInclude or excludePM's reasoning
User's nameIncludePersonalizes responses; required for handoff to human agent if escalated
Past 3 ordersIncludeNeeded to answer order status and return questions accurately
Full order historyExcludeToo long for context window; older orders rarely relevant; increases latency and cost
Credit card last 4ExcludeNo legitimate reason to expose; increases blast radius of any prompt injection attack
Current subscription tierIncludeModel needs this to answer feature availability questions correctly
Internal cost-to-serveExcludeBusiness-sensitive; user should never see this even if they ask
Your move Write the context spec as a table with three columns: element, include/exclude, and your reasoning. This document becomes the source of truth for what goes into the prompt — and it's the thing you'll reference when engineering asks "can we add X to the context?" or security asks "what data does the model see?" Write it before the first line of code, not as a post-launch audit.
03

System prompt collaboration

The system prompt is the most powerful product lever you have — and the one PMs most often hand entirely to engineering. It defines the model's persona, constraints, tone, and operating rules. The PM doesn't have to write it, but the PM has to own what it says.

Prompt elementWho writes itWhat the PM specifies
Persona & rolePM + eng collabWhat the model is, what it's called, what voice/tone it uses, whether it uses first person. This is a brand decision, not a technical one.
Behavioral rulesPM ownsThe "must always" and "must never" list. These are product constraints, not engineering preferences. Engineering implements; PM decides.
Output formatPM specifies, eng implementsResponse length, structure (bullets vs. prose), how to handle uncertainty, when to ask clarifying questions vs. attempt an answer.
Escalation triggersPM ownsExactly when the model hands off to a human, and what it should say when it does. Ambiguous handoffs create the worst user experiences.
Few-shot examplesPM reviewsEngineering often writes these; PM should review them for tone, brand alignment, and whether the examples reflect realistic user inputs (not idealized ones).
Context injection instructionsEng implements, PM specifiesHow the model should use the context it receives — "use the customer's name," "cite the policy document you retrieved," "never mention data you weren't given."
The handoff failure pattern: PM writes a one-line description ("it should be helpful and polite"), engineering writes the actual system prompt, and the PM sees it for the first time in QA. At that point, changing it is a "bug" not a "product decision," and it often doesn't get changed. Review the system prompt as seriously as you'd review a key UI flow — it is the product.
Your move Schedule a "system prompt review" session with engineering before the first external test. Read the current prompt out loud. Ask: does this sound like our brand? Would I be comfortable if a user somehow saw this? Does it cover the edge cases in the Unforgivables section? Does it tell the model what to do when it doesn't know the answer?
04

The Unforgivables section

The most important section of an AI PRD that no traditional PRD has. A list of things the model must never do, regardless of how cleverly the user asks. This is the place where you translate legal, brand, and ethical requirements into explicit model constraints — and where you draw the lines that protect the company.

Write unforgivables as complete sentences the model could follow literally. "Never discuss competitor pricing" is better than "handle competitors appropriately." Ambiguity in the Unforgivables section becomes policy ambiguity in production.
Your move Write your product's Unforgivables section in collaboration with Legal, Brand, and — where relevant — Compliance. Then red-team it: give 20 adversarial prompts designed to elicit each unforgivable to a colleague and have them try to get the model to cross the line. The ones that succeed belong in the system prompt as explicit override instructions, not just in the PRD.
05

Edge cases & failure mode documentation

AI edge cases are not bugs — they are the predicted failure modes of a probabilistic system. Documenting them is the PM's job, not a post-launch surprise. A PRD that doesn't document failure modes will produce a product whose failures feel unpredictable, because nobody planned for them.

Failure modeWhat it looks likeRequired mitigation in PRD
HallucinationModel states a fact with confidence that isn't in its context and isn't trueDefine faithfulness threshold; specify what the model should say when it doesn't know; require source citation for factual claims
Over-refusalModel declines to answer reasonable questions because of overly cautious guardrailsDefine refusal rate target; enumerate categories where refusal is correct vs. incorrect; build override path
Context confusionModel uses the wrong user's data, mixes up session context, or references information it shouldn't haveSpecify context isolation requirements; define what the model should do if context is incomplete or ambiguous
Tone driftModel gradually adopts a different voice over a long conversation, especially after user pushbackSpecify that system prompt instructions persist regardless of user conversation; test for tone stability in multi-turn evals
Prompt injectionUser or retrieved content attempts to override system instructionsSpecify that system prompt takes precedence over user instructions on any constraint; test with adversarial inputs
Confident uncertaintyModel answers with high confidence on topics where it should hedge or escalateDefine escalation triggers by topic category; specify uncertainty language when model is operating near the edge of its knowledge
Your move For each failure mode: write one test case that would produce it and specify the desired behavior. This becomes your red-team test suite. If the team can't write a test case for a failure mode, they probably haven't thought through the mitigation — and the failure will find you in production instead.
06

Full AI PRD structure

Every section below maps to a real product decision. Sections marked AI-specific have no equivalent in a traditional PRD. Sections marked expanded exist in traditional PRDs but need significantly more depth for AI products.

1. Problem & OpportunityStandard
What's in itWhat user problem does this solve, what's the business case, what would a non-AI solution look like and why is AI better for this specific use case.
AI-specific additionExplicitly answer: why does this problem benefit from AI specifically? If the answer is "a rules engine could do most of this," document the marginal value of the AI approach and its marginal cost and risk.
2. Architecture decisionAI-specific
What's in itPrompt engineering vs. RAG vs. fine-tuning vs. custom model — the chosen approach and why. This belongs in the PRD, not just in engineering design docs, because it affects cost, latency, maintenance, and risk.
Key questionsWhat's the expected token cost per task? What model are we starting with and why? Is this a single call or a multi-step pipeline? See the Pipeline Calculator for compound reliability implications.
3. Context specificationAI-specific
What's in itThe full context framework from §2 above — every piece of data that flows into the model's prompt, why it's there, where it comes from, what's excluded, and the staleness rules.
Also includeContext window budget — estimated token usage per element, total budget, and what gets truncated first if the window fills.
4. Persona & system prompt specAI-specific
What's in itThe model's name and role, voice and tone, output format rules, escalation triggers, and the PM's sign-off on the system prompt before it ships. See §3 for the full collaboration framework.
5. The UnforgivablesAI-specific
What's in itThe complete list from §4. Every item must be written as a complete behavioral rule, not a category. Legal and Brand sign-off documented here.
6. User stories & flowsStandard
What's differentEvery flow needs a parallel graceful degradation version: what does the user see when the model is wrong, slow, or refuses? A PRD that only documents the happy path is incomplete for an AI product. Don't write "show an error" — write the specific fallback behavior.
Graceful degradation exampleInstead of: "If RAG retrieval fails, show an error message." Write: "If retrieval fails to find a relevant document, the assistant responds: 'I couldn't find a specific policy on that. Would you like me to connect you to support?' — never a raw error state." Define the fallback copy and the escalation path in the PRD, not in a design review six weeks later.
Latency budgetDefine the Time to First Token (TTFT) SLA directly in this section, not just in the engineering spec. e.g. "The model must stream the first token within 800ms at p95 — if TTFT exceeds this, the UI shows a loading state after 500ms and falls back to a cached default response after 3s." Latency is a UX constraint, not just an infra metric — it belongs in the PRD. See Inference Tradeoffs and AI UX Patterns §1.
7. Edge cases & failure modesAI-specific
What's in itThe failure mode table from §5, plus the red-team test cases. Each failure mode paired with its designed behavior and mitigation.
8. Eval & launch criteriaExpanded
What's in itThe specific metric, threshold, baseline, and eval set definition — not just "the model should perform well." Which quadrant in the Go/No-Go Rubric? What slices must be checked? What's the human-performance baseline?
Golden Set ownershipThe PRD must name who is responsible for building the eval set (the Golden Set) and define the minimum size required for v1 sign-off. e.g. "The ML team will build an initial Golden Set of 200 labeled examples by [date], covering at least 5 user intent categories and including 20% adversarial/edge-case inputs. Launch is gated on this set existing and being reviewed by the PM." Without explicit ownership and a size target, the Golden Set becomes a vague pre-launch dependency nobody owns. See Evaluation Methods §4 for bootstrapping approaches and labeling economics.
Never write"Accuracy should be high" or "quality must be acceptable." These are not eval criteria. They are ways to avoid a difficult conversation that will happen in the sprint before launch instead of the sprint before development.
9. Cost & unit economicsAI-specific
What's in itExpected cost per task, gross margin at target volume, cost optimization plan if margins are tight. Use the ROI Calculator to produce the numbers. A feature that fails on unit economics is a failed product, regardless of quality.
10. Monitoring & data flywheelAI-specific
What's in itWhat signals production generates, who owns them, how they feed back into the eval set. Define implicit and explicit feedback mechanisms (see Data Flywheel). A PRD that doesn't specify signal capture means the model never gets better from real usage.
07

Questions to answer before writing the PRD

If you can't answer these, the PRD isn't ready to write — the product isn't ready to be defined.