The two dimensions that actually matter
Forget the raw accuracy number for a moment. The question that determines whether 87% is good enough is: what happens when it's wrong? That answer has two parts.
mistake affects few people or low stakes
mistake is costly, public, or wide-reaching
user or system can undo it
Low stakes, correctable
Wrong answers are visible, fixable, and don't compound. Users tolerate imperfection when they can easily fix it.
AI email drafts, internal summaries, writing suggestions, search results, content recommendations
Wide impact, but undoable
A wrong answer reaches many people but can be retracted, corrected, or appealed. Speed matters; perfection doesn't.
Content moderation, bulk email personalization, product recommendations at scale, customer-facing chatbots
mistake persists or causes downstream harm
Narrow but sticky mistakes
Wrong answers are hard to undo even if few people are affected. The cost is in the fix, not the scale.
Document classification, code generation for production, automated data transformations, invoice processing
High stakes, hard to undo
A wrong answer is expensive, dangerous, or legally consequential, and can't be easily reversed. No amount of accuracy justifies shipping without hard-coded fallback rules.
Medical triage, automated refunds/financial actions, legal document drafting, safety-critical routing
Threshold table by use case
These are starting-point thresholds, not universal rules. Your specific feature, user base, legal context, and fallback design all change the right number. Use these to anchor the conversation, then adjust.
| Use case type | Suggested threshold | Required before shipping |
|---|---|---|
| AI-assisted drafting (internal) | 75–80% | Easy edit affordance, user sees it's AI-generated, no auto-send |
| Customer-facing suggestions / recommendations | 82–87% | Clear labeling as AI, one-click dismiss, CSAT monitoring from day one |
| Content moderation / classification | 88–92% | Human review queue for borderline cases, appeal mechanism, slice eval by protected attributes |
| Customer support deflection | 85–90% | Human escalation path always available, repeat-contact monitoring, CSAT guardrail |
| Automated data transformation / extraction | 90–94% | Confidence-flagged review queue, downstream error monitoring, rollback capability |
| Agentic pipeline (multi-step) | End-to-end >80% | Use the Pipeline Reliability Calculator — individual step accuracy is not sufficient |
| Financial / transactional automation | 95–98%+ | Hard-coded override rules, human approval above $ threshold, audit log, legal review |
| Medical / safety-critical | Do not ship on accuracy alone | Regulatory approval, clinical validation, explicit disclaimer, mandatory human in the loop |
Interactive rubric
Answer four questions about your feature. Get a shipping stance and a list of what needs to be true before you launch.
Think about one failure — not the aggregate. Does it affect one user, a cohort, or everyone?
Can the user undo it? Can your team patch it within hours? Or does it persist downstream?
An eval on production-representative traffic with known baselines is worth far more than one on curated examples.
How to use this in practice
Before the first sprint: Plot the feature on the 2×2. Document the quadrant and the implied threshold in the PRD or project brief. This is the PM's call, not engineering's.
When the eval result comes back: Compare against the threshold for your quadrant. If you're below it, ask two questions before proposing more model work: (1) can we strengthen the fallback to move to a more lenient quadrant? (2) is our eval representative — could the number look different on harder or more realistic examples?
When leadership pushes to ship anyway: Use the quadrant framing, not the accuracy number. "We're in the high-blast-radius / low-reversibility quadrant and we're at 89% against a 95% threshold" is a clearer conversation than "we're at 89%." The executive question you're answering is: what's the worst case and can we live with it?
The "stop building" case: If the feature is in the red quadrant and the team can't get above 92% after two meaningful iterations, seriously consider whether the fallback design (not the model) is the missing piece. A 90% model with a well-designed human escalation path is often better product than a 97% model with no fallback. See UX for Failures for the design side.
For agentic pipelines: The individual step accuracy is not the right number to assess. Use the Pipeline Reliability Calculator to get the end-to-end rate first, then apply this rubric to that number.