01
Make the case
You are asking for a decision: funding, headcount, a slot on the roadmap, or permission to ship. The pitch that works is not the one that sounds most confident — it is the one where the assumptions are visible, so leadership can argue with the numbers instead of with you.
Six things a decision-maker needs. Anything else is detail you can bring if asked.
The six-part ask
1The decision you are asking for, and by when. "Approve two engineers for a quarter to take this to a limited launch" is a decision. "Invest in AI" is a wish. Name the date the decision expires and what it blocks.
2The problem in their units. Support cost per ticket, hours per week, conversion, churn. If you cannot state the problem in a number leadership already tracks, you are asking them to care about a new metric and fund it at the same time.
3Why AI, and not rules. The question is coming. Answer it before it is asked — see §2.
4Expected value, with the assumptions showing. Not "this will save $2M." Instead: "At 40,000 tickets a month, a 30% deflection rate at 85% quality saves roughly $X. The two assumptions that move this most are deflection rate and what an escalated ticket costs." Put the sensitive assumptions on the slide. Use the ROI calculator to get the range, and quote the range, not the point estimate.
5The risk, said out loud. The top two or three failure modes, the rate you expect, and what catches each one. A case with no risks reads as naive and invites someone else to find them for you. State where the feature sits on blast radius and reversibility (see the Go / No-Go rubric) and what the fallback is.
6What you need, and what you will report. People, data access, and time — plus the checkpoint: "In six weeks I will bring back the eval against the rules baseline, cost per completed task, and a go/no-go recommendation." A named checkpoint converts a bet into a staged decision, which is far easier to approve.
The number that kills a case later: a value estimate with no baseline. If you cannot say what today's process achieves and costs, the claimed improvement is unfalsifiable — and the first person to ask "compared to what?" wins the room. Get the incumbent and heuristic numbers first (see
Baselines).
Your move
Write the ask as one sentence before you build the deck: "I'm asking for [decision] by [date], to [outcome in their units], at [cost], with [risk] mitigated by [fallback]." If that sentence does not hold together, the deck will not fix it — and you have just saved yourself a week.
02
Why AI, and not just rules?
This question arrives from both directions. A skeptical engineering leader asks why you cannot just write the logic. An enthusiastic executive asks why you are not using AI for everything. The same evidence answers both, and having it ready is the difference between looking rigorous and looking like you followed a trend.
| Choose rules when | AI earns its cost when |
| The logic is stable and someone can write it down |
The variety outruns what anyone can enumerate — and the tail is where the value is |
| Every decision must be explained and reproduced exactly |
"Good" is a judgment people make consistently but cannot fully specify |
| The inputs are structured and predictable |
The inputs are unstructured — text, speech, images, documents |
| Volume is low enough that people can handle the exceptions |
The pattern shifts over time, so rules go stale faster than anyone can maintain them |
| A wrong answer is unacceptable and the cases can be enumerated |
Being right most of the time, with a human path for the rest, is genuinely valuable |
How to say it in the room
When asked "why not just write rules?": "We built the rules version — it handles about 70% of cases and breaks on the rest, because the variety is in the wording, not the structure. The model covers that tail. We are keeping rules for the cases we have to guarantee, and the model handles what rules cannot enumerate."
When asked "why aren't we using AI here too?": "That one is a rules problem. The logic is stable, it has to be auditable, and a model would add cost and latency to get the same answer less predictably. We would be paying for uncertainty we do not need."
The strongest version of the argument is a number: the rules baseline scored X, the model scored Y, and the gap in business terms is Z. That sentence ends the debate faster than any framing — which is why establishing a heuristic baseline is worth doing before the meeting, not after.
The honest answer is usually "both": rules for the cases you must guarantee, a model for the long tail, and a human for what neither should decide alone.
Your move
If you cannot clearly state why rules fall short for this problem, that is not a presentation gap — it is a signal worth taking seriously. Run
Does this actually need AI? before you build the case. Coming back with "we checked, and rules are the better answer here" builds far more credibility than shipping a model that a stable rule would have beaten.
03
The n=1 problem
An executive tests a prompt, gets a bad answer, and concludes the product is broken. They're not wrong to raise it — but they're using the wrong mental model to interpret it. Your job is to give them the right one without making them feel corrected.
AI failures are distributed, not discrete. A bug in traditional software fails the same way every time and can be patched permanently. An AI failure is a sample from a distribution — fixing it for that exact input doesn't mean it's fixed everywhere, and it working for another input doesn't mean it's fixed at all.
The executive mental model: software either works or it doesn't. A bug gets filed, a fix gets shipped, the problem goes away. The AI mental model: the product operates at a quality level across a distribution of inputs. Some inputs get bad outputs. The question isn't "is it broken?" — it's "what percentage of inputs get bad outputs, and is that acceptable for this use case?"
The reframe that usually lands: "Think of it less like a bug report and more like a quality audit. We measure accuracy across thousands of inputs, not just one. The question is: what does our distribution look like, and did this example reveal a gap in a slice we care about?"
Before the conversation
Have your eval numbers ready, not to recite them, but to reference them. "That failure mode appears in about X% of cases like this one" is a completely different conversation than "I'll look into it." The first anchors the exec in the distribution; the second confirms their suspicion that something is fundamentally wrong.
04
Response scripts
The goal of each script is the same: acknowledge the failure, reframe it from "bug" to "distribution sample," and redirect toward what actually matters. Every script has a "say this, not that" structure.
Scenario 01 — the angry test
"I just tested it and it gave a completely wrong answer. This is not ready."
Don't defend the output or get into why that specific answer was wrong. Don't promise a fix for that exact prompt. Acknowledge it, reframe to the distribution, and give them a way to evaluate quality that actually scales.
"You're right to flag it — that's a bad output. What I want to make sure we're doing is evaluating this at scale, not just on single examples, because AI quality works differently than traditional software bugs. We measure across thousands of inputs. What you saw is real and worth investigating, but the decision should be based on what the distribution looks like, not this one case. Let me pull the eval data and show you what percentage of similar inputs fail, and whether that's within our shipping threshold."
Scenario 02 — the comparison
"I tested ChatGPT on the same prompt and it did better. Why are we building this?"
This one is about evaluation context, not about defending your model. The exec is comparing single outputs from two different systems with different prompts, tuning, and use cases. Redirect to what a real comparison looks like.
"That's a fair instinct to compare, but a single prompt isn't how you measure model quality — it's like comparing two cars by one test drive on different roads. The right comparison is running both systems on the same set of real customer inputs, scored the same way. We do that benchmarking as part of our eval process. On our task distribution, our system [performs / is competitive / has specific advantages in X] because it's tuned for our context and our data. Want me to show you what that looks like?"
Scenario 03 — the board pressure
"The board is asking why we're not shipping this faster. Can we just get it out?"
This is a blast radius and reversibility conversation, not an accuracy conversation. Bring it back to risk, not capability. If you've done the Go/No-Go rubric, reference it — it becomes your cover and the exec's cover.
"I want to ship as fast as we can too. The constraint isn't capability — it's that this feature sits in a [high blast radius / low reversibility] zone. If we get it wrong at scale, the cost to fix it is [describe the consequence]. Our eval shows we're at X%, and our threshold for this risk category is Y%. We're [Z weeks] from there. Shipping early saves [time] but risks [specific consequence]. I'd rather give you that trade-off explicitly than discover it after launch."
Scenario 04 — the regression
"It worked fine before. You shipped something and broke it."
Don't immediately accept that the model regressed. This might be drift (the world changed, not the model), a new input pattern, or a real regression. The key distinction — which you need to diagnose before the conversation — is whether something you shipped caused this or something in the environment did.
"The first question I need to answer is whether this is a model regression — something we changed — or a distribution shift — something about how the inputs changed. Both look the same on the surface. I'm running a regression analysis against our holdout set now. If we caused it, I'll have a fix timeline. If it's distribution shift, the fix is different — we retrain or update on the new pattern. Either way I'll have an answer within [timeline]."
What all four scripts have in common
Each one buys you time to show data, reframes from discrete bug to distribution, and ends with a concrete next step. The pattern: acknowledge → reframe → redirect to the right measure → commit to a timeline. Never leave an exec conversation without a named next step and a specific date.
05
Board deck framing
F1 scores and precision/recall don't belong in a board deck. Not because boards can't understand them — they can — but because they don't map to what a board actually needs to evaluate: business risk, competitive position, and investment return. Here's how to translate.
| Don't say (ML metric) | Say instead (business frame) | Why it lands better |
| F1 = 0.87 | The model handles 87 out of 100 cases correctly without human review. We're targeting 92% before broad rollout. | Gives a concrete picture of the human-to-AI handoff ratio and shows you have a defined target. |
| Precision = 94%, Recall = 79% | When it acts, it's almost always right. The risk is missed cases, not wrong actions — and we route those to humans. | Translates into the business risk they care about: false alarms vs. missed opportunities. |
| Hallucination rate = 3.2% | About 1 in 30 responses contains an error the user might not catch. We have guardrails and a correction mechanism, and we're monitoring this weekly. | Scale + mitigation in one sentence. Shows you know the risk and have a plan. |
| Latency p95 = 4.2s | 95% of responses arrive in under 4 seconds — within the range where user research shows people wait without abandoning. We alert when this degrades. | Anchors in user behavior, not infrastructure metrics. |
| Cost per inference = $0.012 | The AI handles X,000 tasks/month at a cost of $Y — roughly $Z per task, versus $W per task with humans. Margin at this volume is [X]%. | This is the number CFOs actually need. See the ROI Calculator. |
| We're using GPT-4 / Claude / Gemini | We abstract our model vendor to avoid lock-in. We benchmark across providers quarterly and can switch within [weeks] if pricing or quality shifts. | Answers the vendor risk question before it gets asked. |
The three slides that cover 90% of board AI questions
Slide 1 — Quality snapshot: one chart showing accuracy trend over time, the target threshold, and current slices. Not the methodology — the number, the trend, and the target. Add one sentence on what "wrong" looks like and who catches it.
Slide 2 — Unit economics: cost per task vs. human cost per task, gross margin at current volume, and the margin at projected scale. This is the slide that gets the CFO on your side. If the numbers aren't there yet, show the path and the volume where they flip positive.
Slide 3 — Risk register: the top three failure modes, what the current rate is, and what the mitigation is. Boards want to know you've thought about what can go wrong — not that nothing can. The absence of a risk register reads as naivety, not confidence.
Your move
Write the board summary before you're asked for it. When a VP asks "how do I explain this to the board?", having three clean slides ready is the difference between being seen as a strategic partner and being seen as someone who needs to be managed up. Build the translation layer into your regular reporting rhythm so you're never caught cold.
06
AI SLA template
A traditional SLA promises 99.9% uptime. An AI SLA does something different: it sets expectations about quality under uncertainty, commits to safety guardrails, and disclaims the guarantee of determinism — without sounding like a legal disclaimer nobody reads.
What an AI SLA is not: a promise that every output will be correct. What it is: a documented, signed-off commitment about quality levels, what happens when the system fails, who is responsible for what, and what the human fallback looks like. It protects both the user and the team.
AI SLA — template clauses
1The system operates at a measured accuracy of [X%] on [task description], evaluated on a representative sample of [N] real inputs, with a human baseline of [Y%] and an incumbent baseline of [Z%]. This measurement is updated [quarterly / on every major model update].
2Non-determinism disclosure: The system may produce different outputs for identical inputs across sessions. This is expected behavior, not a defect. Quality commitments are made at the distribution level (across many inputs), not at the individual output level.
3Human fallback: All outputs in [category / above / below confidence threshold] are [reviewed by a human before action / flagged for human review / surfaced with a confidence indicator]. A human escalation path is always available at [describe the mechanism].
4Unforgivables (hard limits): The system will never [list 2–4 specific actions or outputs that are unconditionally off-limits]. These are enforced by hard-coded rules that operate independently of the model, and are tested on every deployment.
5Monitoring commitment: We monitor [key metrics: error rate, latency, hallucination rate, CSAT] in real time. Alerts fire when any metric exceeds [threshold]. A responsible team member is on call to respond within [SLA window].
6Incident response: When a material quality degradation is detected (defined as [X% drop in accuracy / Y% increase in error rate]), we will [describe response: rollback, human-only mode, stakeholder notification] within [timeframe].
7User recourse: Any user can [report an incorrect output / request human review / override the AI decision] at any time. Feedback is logged and reviewed [weekly / monthly] and feeds directly into eval set improvement.
Your move
Get this signed off internally before you need it externally. The moment a customer or exec asks "what's your SLA on AI quality?" you want to hand them a document, not improvise an answer. The act of writing it also forces the right internal conversations: what are the hard limits? who owns the human fallback? what triggers an incident? those are conversations worth having before launch, not after a failure.
07
Getting ahead of it
The best executive conversation is the one that never happens — because leadership already has the mental model before the angry Slack message arrives. Three things that prevent the crisis mode conversation:
Run a controlled "break the bot" session early
Before the CEO finds the failure on their own, find it with them. Invite the exec team to a structured red-teaming session: "We're going to deliberately try to make this fail." When they find edge cases themselves, in a controlled setting, two things happen: (1) they feel ownership over the product's known limitations, and (2) those edge cases become part of the test suite. The session reframes "it has failure modes" from a PM admission to a shared team finding.
Set a "quality floor" expectation explicitly, once
Early in the product's life, have one explicit conversation with your exec: "This system will be right X% of the time at launch. Here's what wrong looks like and here's who catches it. We're targeting Y% by [date]." This conversation, had once and documented, means every future n=1 test failure gets evaluated against an agreed baseline — not treated as evidence the whole thing is broken.
Build a monthly quality brief for non-technical stakeholders
One page, once a month: accuracy trend, cost per task, the top three failure modes and their rates, one thing that improved, one thing that didn't. Not a metrics dump — a narrative. Leadership that gets this monthly develops an accurate intuition for AI quality over time, which means fewer panicked escalations and better funding conversations. It also creates a paper trail that protects you when things go wrong.
The bottom line
You cannot eliminate AI failures. You can control whether they feel like surprises or expected variance. The difference is almost entirely about communication, not capability — and communication is entirely the PM's job.