Latency patterns
AI responses are slow relative to user expectations set by deterministic software. The solution isn't always to make the model faster — it's to design the experience so the wait doesn't feel like waiting.
Token-by-token rendering
- Show output as it generates — first token often arrives in <1s even when full response takes 8s
- Dramatically lowers perceived latency
- User starts reading before generation finishes
- Works best: conversational responses, long-form drafts, explanations
Pre-generate or async
- Trigger generation before the user asks — on page load, on hover, on prior step completion
- Result appears "instantly" because it was already computed
- Works best: suggestions, completions, next-step recommendations, summaries
- Risk: wasted compute if user doesn't trigger the feature
| Pattern | How it works | Best for |
|---|---|---|
| Optimistic UI | Show a placeholder or skeleton of the expected output immediately; replace with real output when ready | Structured outputs where the shape is known before content |
| Progressive disclosure | Show a summary first (fast), offer "expand" for the full output (slow) | Long documents, detailed analysis, research summaries |
| Loading narrative | Show what the AI is doing step by step ("Searching knowledge base… Drafting response…") rather than a spinner | Multi-step agents, RAG pipelines — makes the wait feel productive |
| Speculative generation | Pre-generate likely next responses in the background based on conversation flow | Chatbots and assistants with predictable conversation trees |
| Smaller model first | Return a fast, cheaper model's answer immediately; offer to "improve with more thinking" for an upgraded result | Any use case where 80% of users need the quick answer |
Friction as a feature
Not every AI action should be instant and automatic. Sometimes adding a deliberate pause — a review step, a confirmation, a visible intermediate state — is the right product design. Friction, used intentionally, builds trust and catches errors before they compound.
Copilot acceptance states
Show AI-generated content in a visually distinct "pending" state before the user commits to it. GitHub Copilot's inline suggestions, Notion AI's shaded blocks, and Gmail's Smart Compose all use this pattern. The friction is the single keystroke or click to accept — small enough not to break flow, deliberate enough to prevent mindless acceptance.
Confirmation before irreversible actions
When an AI agent is about to take an action that can't be undone — sending an email, submitting a form, making a purchase, deleting data — a confirmation step is not bureaucracy. It's the product acknowledging that the AI could be wrong and giving the user one last chance to check. The cost is one click; the benefit is a user who stays in control and trusts the system.
Low-stakes, high-reversibility actions
If the action is trivially reversible (auto-categorizing an email, tagging a note, suggesting a calendar slot) and the cost of being wrong is low, friction creates more annoyance than protection. The pattern here is to act optimistically and surface an undo, not to block with a confirmation. Every unnecessary confirmation trains users to click through without reading.
Disambiguation patterns
Sometimes the model doesn't have enough information to act confidently. The question is: what does the product do? Silently guessing is usually wrong. Asking the user in a clunky dialog is friction that kills flow. The patterns below sit between those two failure modes.
Interpret and show your work
Act on the most likely interpretation, but surface it explicitly so the user can correct it without extra steps. "I'm treating this as a refund request for your March 12 order — let me know if that's wrong." This keeps flow intact while catching misunderstandings before they compound downstream.
Offer structured options, not open questions
When the model needs clarification, give the user 2–3 concrete options to pick from rather than asking an open-ended question. "Do you want a short summary, a detailed analysis, or bullet points?" is faster to answer and less cognitively demanding than "What kind of output do you want?"
One question, at the right moment
For genuinely high-stakes or high-effort generation (a long document, a complex workflow, an agentic task), asking one scoping question before starting is better than producing the wrong thing and iterating. The rule: one question maximum, at the start, only when the ambiguity would materially change the output. Never interrupt mid-generation to ask.
Graceful degradation
What does your product look like when the AI fails? Not eventually, but right now, on this request, for this user. Designing the failure state is as important as designing the success state — and it's almost always done last.
| Failure type | What the user sees | What to design |
|---|---|---|
| API timeout | Blank screen, spinner that never resolves, generic error | Show a fallback state with a retry option. "This is taking longer than expected — try again or continue without AI assistance." Never leave the user staring at a spinner with no recourse. |
| Low-confidence output | A confidently-phrased wrong answer, indistinguishable from a correct one | Surface uncertainty visually. Hedging language, a confidence indicator, or a "flag this response" affordance. Only when the model is actually calibrated — see Confidence & Calibration. |
| Out-of-scope request | A refusal, an off-topic response, or a confabulated answer | Design an "I can't help with this" state. Not a generic error — a specific redirect: "I can't help with that, but here are 3 things I can do." Route to a human or relevant resource when available. |
| Safety block / guardrail trigger | An opaque "I can't do that" with no explanation | Explain what happened without revealing the guardrail details. "I'm not able to help with that type of request. Here's what I can help with instead." Never leave the user with no path forward. |
| Wrong answer the user caught | A bad output — factually wrong, tonally off, structurally broken | Make correction trivially easy. Inline editing, one-tap regeneration, "try a different approach." Measure override rate as a product metric. If users can't easily fix it, they'll stop trusting it — not just stop using the feature. |
Building trust over time
Trust with AI products isn't binary — it's a spectrum that users move along based on their experience. New users need more scaffolding. Established users need less friction. Designing for this arc is different from designing for a single interaction.
Show the reasoning, not just the result
New users don't know what the AI is capable of or where it fails. Showing why the AI made a decision ("I'm suggesting this because your last 3 orders were from this vendor") helps users build an accurate mental model of the system's strengths and limits — faster than they'd develop it through trial and error.
Reduce scaffolding as confidence grows
Confirmations that were helpful in week one become friction in week six. Progressive trust-building means surfacing fewer intermediate steps for users who've demonstrated they understand and check the outputs. This can be explicit (let users set their trust level) or implicit (reduce confirmation frequency after N accepted outputs).
Overtrust and automation bias
Users who've had a long run of correct outputs often stop checking. This is the moment a model degradation or distribution shift becomes most dangerous — the users who trusted it most will catch it least. Build in periodic "spot check" moments, visible accuracy cues, and make it easy to reactivate review mode. Don't let trust become complacency.
AI UX design checklist
Run this before any AI feature ships — not during QA, but during design review. If the answer to any of these is "we haven't decided yet," that's a product gap.
| # | Question |
|---|---|
| 1 | What's the TTFT target, and have we designed the experience for the p95 case, not just the average? |
| 2 | Is every AI-generated output visually distinguishable from user-generated content? |
| 3 | Can a user override or edit any AI output in ≤2 steps? |
| 4 | What does the user see when the API times out? |
| 5 | What does the user see when the model returns a low-confidence output? |
| 6 | What does the user see when the model refuses or goes out of scope? |
| 7 | For every irreversible action: is there a confirmation step that describes the specific action? |
| 8 | For every ambiguous input type: is the disambiguation strategy documented and designed? |
| 9 | Does the first-session experience differ from the established-user experience? |
| 10 | Is there a mechanism to detect and respond to user overtrust or automation bias? |