Baselines: is this number any good?
"The model is 85% accurate." Good? Terrible? You cannot answer that without a baseline — and picking which baseline is a product judgment, not a technical one. A model that beats random but loses to the ten-line rule your team already runs is not a win.
Zero-rule / majority class
Random
Simple heuristic
Human performance
The incumbent — previous model or current system
Baselines differ enormously in what they cost to produce — which is why the cheap ones get reported and the useful ones get skipped. Knowing the effort ahead of time lets you ask for the right one instead of being told it's not feasible.
| Baseline | Where the data comes from | Effort | Watch out for |
|---|---|---|---|
| Zero-rule | Computed directly from your eval set's label distribution. No new data, no new system — just count which class is most common. | Minutes. Effectively free. | Nothing. There is no excuse for not having this one, and its absence usually means nobody checked class balance at all. |
| Random | Same label distribution, sampled randomly. Also computed, not collected. | Minutes. Effectively free. | It's so weak that beating it proves almost nothing — don't let it stand in for a real comparison. |
| Simple heuristic | If a rules-based system already runs in production, you have it — just re-run it over the eval set. If not, someone writes the rules and runs them. | Hours to a day if it needs writing. Nearly free if a rules system already exists. | Often dismissed as "not worth the time." It usually takes less time than the meeting about whether to do it — and it's the comparison most likely to change your decision. |
| Human performance | Sample a few hundred examples from the eval set and have two or more people label them independently. Their agreement with ground truth gives you human accuracy; their agreement with each other gives you the practical ceiling. | Days, and real people-hours. The expensive one. | Do it once on a stable sample and reuse it — it doesn't need re-running every cycle. Skipping the two-rater part loses the ceiling number, which is the more useful half. |
| Incumbent | Two valid routes: re-run the current system over the same eval set, or run the new model in shadow mode on live traffic (it scores every request but its output isn't shown) and compare on identical inputs. | Low to re-run offline. Moderate to set up shadow mode — but shadow mode is reusable infrastructure. | The big one — see the rule below. |
Brand-new capability, no incumbent, no labels, no historical data. This is common and it's not a dead end — it just changes what you can claim.
| Option | What it gives you |
|---|---|
| Hand-label a small golden set | A few hundred carefully labeled examples is usually enough to compute zero-rule, random, and human baselines. This is the standard starting move — and it doubles as your eval set (see Evaluation Methods §4). |
| Write the throwaway heuristic anyway | Even a deliberately crude rule gives you a floor to beat. If the model can't clear it, you've learned something important cheaply. |
| Ship to a small slice and build labels from outcomes | Real user behavior — clicks, corrections, escalations, refunds — becomes your ground truth over time. Slower, but produces labels that reflect actual production reality rather than a curated set. |
| Use human performance as the only anchor | When there's genuinely nothing to compare against, "how well does a trained person do this?" is a legitimate and defensible bar on its own. |
Slice-based evaluation
An aggregate metric is an average, and averages hide their worst cases. A model at 92% overall can be at 97% for your largest segment and 61% for the one you just signed a contract with. The slices that matter are business segments, not statistical ones — and you're the person who knows which ones are commercially load-bearing.
Same model, same data. The 92% is real — and it's also the least useful number on this page. New users at 61% is an onboarding and retention problem. Enterprise at 78% is a churn and renewal problem. Non-English at 58% may be a compliance or market-entry problem.
A B2B tool that reads invoices and extracts fields automatically
Nothing about the model changed in that story. What changed was which number the team was accountable for — and that came from a PM asking who sat inside the average.
| Slice type | Examples | Why a PM cares |
|---|---|---|
| Customer value | Enterprise vs. SMB, paid vs. free, top-decile accounts | A failure concentrated in your highest-revenue segment is a revenue problem, not a metrics problem. |
| Lifecycle stage | First session, first week, established, reactivated | New-user failures hit activation and retention hardest — and those users have the least patience for a bad first experience. |
| Input characteristics | Short vs. long queries, image quality, language, rare categories | Where "the long tail" actually lives. Usually the gap between demo-quality and production-quality. |
| Channel / platform | Mobile vs. desktop, API vs. UI, region | Often reveals an integration or context-window problem rather than a model problem. |
| Protected attributes | Where legally and ethically relevant | Fairness obligations (see S-05) — and a category of risk that does not show up in any aggregate. |
Confidence & calibration
Most models output a confidence score alongside their answer. The question that matters: when the model says it's 90% sure, is it actually right about 90% of the time? That's calibration — and it's separate from accuracy.
Confidence means something
Of all the predictions made at 90% confidence, roughly 90% turn out correct. You can set a threshold and trust it — auto-approve above it, route to a human below it, and know roughly what you're getting.
Confidence is decoration
The model says 95% on things it gets wrong half the time (overconfident) or hedges at 60% on things it nails (underconfident). Any threshold you set is arbitrary, and showing that number to users is actively misleading.
Calibration is what makes confidence-based routing possible, and that pattern is one of the most practical tools a PM has: handle the confident cases automatically, send the uncertain ones to a human. It converts a model that's 85% accurate overall into a system that's ~99% accurate on the 70% of cases it handles alone — and escalates the rest. But it only works if the confidence score is honest.
The offline–online gap
The model scored well in the eval and then underperformed in production. This is the norm, not the exception, and the reasons are mostly predictable.
| Cause | What's happening | What to do |
|---|---|---|
| Data mismatch | Real traffic doesn't look like the eval set — messier, shorter, more typos, more edge cases (see survivorship bias in Evaluation Methods §5). | Rebuild the eval set by sampling from actual production traffic, not curated examples. |
| Feedback loops | The model's own outputs change user behavior, which changes future inputs. Recommenders are especially prone to this — you only see clicks on what you showed. | Reserve a small slice of random/exploratory traffic so you can measure what you'd otherwise never observe. |
| Latency & timeouts | Offline, every prediction completes. Live, some requests time out, get truncated, or fall back — and those failures aren't in the eval score. | Measure end-to-end success including failures, not just model accuracy on completed calls. See Inference Tradeoffs. |
| Integration gaps | The model is fine; the retrieval, the prompt assembly, the post-processing, or the UI is losing the value. | Evaluate the system end to end, not the model in isolation. Users experience the pipeline, not the model. |
| Population shift | Your users changed — new market, new campaign, new segment, seasonality. | Monitoring (§6), not a better eval set. The eval wasn't wrong; the world moved. |
Metrics by business type
The same model quality means completely different things depending on how the business makes money. Below: the business metric each type actually runs on, the product-layer metric that bridges to it, and — critically — the guardrail that stops you optimizing yourself into a hole.
E-commerce / retail
Marketplace / two-sided
Subscription / SaaS
Customer support
Advertising / media
B2B SaaS & copilots
Monitoring in production
Models degrade without anyone touching them. The distinction that matters most — and the one PMs most often miss — is between the model changed and the world changed. They look identical on a dashboard and need completely different responses.
The inputs changed
The distribution of what users send has shifted. The model is doing exactly what it always did — it's just seeing different things.
The right answer changed
Same inputs, but what counts as correct has moved. The model is now confidently wrong by an old definition.
The model changed the world
The model's own outputs shaped user behavior, which now shapes its inputs. Hardest to detect because everything looks internally consistent.
| Cadence | Watch | Because |
|---|---|---|
| Alert (minutes) | Error rate, latency p95/p99, traffic volume, model availability, output-format failures | Something is broken right now and users are feeling it. |
| Daily | Confidence distribution shifts, input-length and input-mix changes, fallback rate, guardrail metrics | Early signals of drift — visible before quality metrics move. |
| Weekly | Slice performance, override/edit rate, CSAT, escalation rate | Quality trends that need human eyes and context, not a pager. |
| Monthly / quarterly | Business-layer metrics, cohort retention, cost per resolved task, incumbent comparison | Did it actually deliver the business outcome you promised when you shipped it? |
Unit economics & ROI
An accurate model that destroys your gross margin is a failed product. Business fit isn't just about whether the model works — it's whether it works at a price the business can sustain. LLM API calls and ML infrastructure are notoriously expensive at scale, and the unit cost compounds fast in agentic and multi-call workflows.
Cost per inference vs. value generated
Every AI call has a marginal cost — API token spend, compute, retrieval, orchestration overhead. Compare that against the marginal value it generates: revenue per recommendation, human time saved per ticket, error cost avoided per document review. If cost per inference is rising faster than value per inference, the unit economics are degrading even if quality is improving. See C-01 and C-02 in the field guide.
Agentic cost compounding
A single-call cost that looks trivial becomes significant when an agent chains 10–30 calls per user task, retries on failure, or calls a judge model on every output. The correct unit is cost per completed task, not cost per token. Budget for retries, fallbacks, and the judge — and alert when task-level cost drifts, not just when individual calls get expensive.
Margin thresholds before you ship
Define the gross-margin floor before launch, not after the CFO asks. A support automation that deflects a $12 ticket while costing $9 in inference per session has a 25% gross margin — thin but viable. At $11, it's destroying value. Neither number is visible unless you're measuring it. Run the model, multiply by your p50 session cost, and decide whether the unit economics work at your projected volume.
| Optimization lever | What it changes | What you give up |
|---|---|---|
| Smaller / cheaper model | Cost per call drops significantly | Quality — validate on your own eval set before assuming it holds |
| Quantization | Memory and inference cost, roughly halved at INT8 | Some quality — see Inference Tradeoffs §4 for what to check |
| Prompt caching | Prefill cost on repeated prefixes (e.g. long system prompts) | Nearly nothing, when prompts share structure |
| Confidence routing | Expensive model only handles the uncertain cases; cheap model or rules handle the rest | Engineering complexity; requires calibration (§3) |
| Batching async work | Throughput up, cost per token down | Latency — only viable for non-interactive workflows |
UX for failures & graceful degradation
No AI feature is right 100% of the time. The product question isn't whether failures happen — it's whether users can tell when they're happening, recover cleanly, and still trust the system afterward. Designing for the failure case is as important as designing for the success case.
-
Show uncertainty
Surface confidence, don't hide it. When the model is less certain, the UI should communicate that — not with raw probability scores, but with language or design that calibrates the user's trust. A confidently wrong answer with no hedging erodes trust permanently; a hedged answer that turns out to be uncertain but honest preserves it. e.g. "Here's a suggested response — review before sending" vs. "This is your response" even when the model is equally uncertain in both cases.
-
Easy override
Make correction effortless and visible. The edit/override rate (see B-06) tells you how often users disagree — but only if you make it easy enough that they actually try. If correcting the AI is harder than doing the task from scratch, users will abandon rather than fix, and you'll undercount disagreement while miscounting success. e.g. Inline editing on AI-generated drafts, one-tap rejection with no explanation required, "not what I needed" escape hatch that doesn't restart the whole flow.
-
Fallback UI
Design the "I can't help with this" state. Low-confidence or out-of-scope requests need a graceful landing — not an error screen and not a confidently wrong answer. Routing to a human (with context preserved), surfacing relevant documentation, or simply saying "I'm not sure" with a clear next step is a better product experience than pushing through. e.g. Below a confidence threshold, the UI switches to "Here are 3 related articles" or "Connect with a specialist" rather than showing a guess.
-
Preserve user state
When the AI fails, the user shouldn't lose their work. A session where the AI gives a bad answer and the user has to start over from scratch is a double failure: wrong output plus lost time. Failures should be recoverable — the user's input, context, and progress should survive a model error. e.g. If an AI code suggestion breaks a build, the IDE should let you one-click revert to before the suggestion was applied.
-
Calibrate trust over time
New users need more transparency; established users need less friction. Early in a relationship, showing more of the model's reasoning and making overrides prominent builds appropriate trust. As users learn the system's strengths and weaknesses, they can be trusted to need fewer guardrails — and adding them anyway creates friction that drives abandonment. e.g. A "why did you suggest this?" affordance matters more in week one than week twelve.
Building the data flywheel
A model trained once on static data degrades as the world changes. The products that compound over time are the ones where what happens in production makes the next model better — automatically. A PM's job is to design that feedback loop, because it doesn't emerge on its own.
What users do, not what they say
Behavioral signals that reveal model quality without asking the user for anything. Higher signal-to-noise, but requires interpretation — a user ignoring a suggestion could mean it was wrong, or just that they were busy.
What users tell you directly
Thumbs up/down, star ratings, flags, "this was wrong" buttons. Lower volume but unambiguous when it happens. Requires visible affordances and enough user motivation to engage — which means it captures the strong opinions and misses the mild dissatisfaction.
Whether the downstream goal happened
The strongest signal, but the hardest to attribute. Did the recommended product get purchased? Did the drafted email get sent without edits? Did the classified ticket resolve correctly? Lagged and noisy, but the closest to ground truth on whether the AI actually helped.
Turning raw signal into training data
Signal alone isn't training data — it needs to be cleaned, deduplicated, and often human-reviewed before it's safe to feed back in. Design the pipeline from the start, not after you have a backlog of uncleaned logs that nobody knows how to use.
Generative AI & RAG metrics
The rest of this page leans on classical predictive-ML language (accuracy, classes, zero-rule) because the frameworks transfer. But if your product is a generative AI or RAG system, the metrics shift — and some don't exist in the classical vocabulary at all. This section is the bridge: which field-guide metrics apply, which are specific to GenAI, and how to evaluate the retrieval and generation layers separately.
Retrieval layer (RAG)
Evaluated separately from the generator — because generation quality is capped by what was retrieved. A generation that's faithful to bad context is still a bad answer.
Production add: latency of the retrieval step separately from generation — if retrieval is the bottleneck, that's an index problem, not a model problem.
Generation layer
Evaluated against the retrieved context (faithfulness) and against the user's actual need (helpfulness). These are distinct — an answer can be faithful to bad context, or helpful but hallucinated.
GenAI-specific: tone consistency and brand safety — does every output sound like the product, and does it avoid language the brand can't stand behind? Usually checked by an LLM judge (see Evaluation Methods §6).
System-level (both layers)
What the user actually experiences — the combined output of retrieval + generation, measured at interaction-level rather than component-level.
Production add: context-window utilization — are you paying for tokens the model ignores? Wasted context is wasted money and sometimes degrades quality (long, irrelevant context can distract attention away from the relevant parts).
Safety & guardrails
Generative outputs introduce risks that don't exist in classifiers — because the output space is unbounded. Guardrails need to be measured, not just assumed.
Watch both directions: too much refusal is a usability problem; too little is a safety problem. Refusal rate needs a target range, not just a floor — see S-03.
Worked example: support deflection, end to end
One case, traced all the way through — baseline, slices, business connection, and the trap. The scenario: a B2B SaaS company ships an AI assistant to answer customer support tickets before they reach a human. Leadership approved it to reduce support cost.
Establish the baselines
The team reports the assistant resolves 68% of tickets without escalation. Before celebrating, we get the comparisons:Zero-rule: always escalate → 0% deflection (the floor).
Simple heuristic: existing keyword-based help-article suggester → 31% deflection.
Incumbent: nothing — this is the first AI system, so the heuristic is the incumbent.
Human: agents resolve 94% without further escalation.
Cut it by business slices
The PM defined the slices up front: customer tier, ticket category, and account lifecycle stage.SMB customers: 74% deflection.
Enterprise customers: 41% deflection.
Billing tickets: 81%.
Integration/API tickets: 38%.
Accounts in first 30 days: 45%.
Connect to the business layer
Support is a cost center here, so the business metric is cost per ticket, and the bridge metric is deflection rate. The named guardrail, set before launch: CSAT must not drop, and repeat-contact rate must not rise.Month one results: deflection 68%, support cost per ticket down 22%. CSAT down 4 points. Repeat-contact rate up 9%.
Decide — as a PM, not as a scientist
The data supports a targeted response rather than a rollback:1. Route enterprise and first-30-day accounts straight to humans. Sacrifices deflection where deflection is least valuable and most risky.
2. Disable the assistant for integration/API tickets until quality improves on that category.
3. Add confidence-based routing (§3) — escalate below threshold instead of attempting an answer.
4. Re-baseline: measure deflection net of repeat contacts, not gross.
Questions worth asking your team
Twelve questions, roughly in the order this page raises them.
- →What are we comparing this number against — and how does it do against a simple heuristic?
- →Which slices did we cut this by, and who's in the worst one?
- →Is the model calibrated — and are we showing or acting on its confidence?
- →How far did production performance land from the offline eval, and do we know why?
- →What business metric is this supposed to move, and what's the guardrail we won't trade away?
- →If the metric moved and we shipped nothing — did the model change, or did our users change?
- →Who reviews the business-layer numbers monthly, not just the alerts?
- →Are we measuring the outcome net of rework — or just the part that looks good?
- →What's our cost per completed task, and does the unit economics work at our projected volume?
- →What does a user see when the AI is wrong — can they tell, and can they recover without starting over?
- →What signal does each AI interaction generate, who owns it, and where does it go — is it feeding the next model?
- →For GenAI/RAG: is this a retrieval failure or a generation failure — and would we stand behind every output publicly?