For PMs building RAG assistants, agents, copilots, and chatbots. Every metric here is connected to what the user does and what the business gets — not just to a benchmark.
Each kit is a workflow: the decisions at every stage, with the metrics, tools, and playbooks that answer them.
Confident answers that aren't in the source — or are in the source but don't answer the question.
Open kit →Every step can fail, and failures compound: five steps at 90% complete only 59% of tasks.
Open kit →Suggestions users accept and then quietly rewrite — usage looks healthy while time saved is zero.
Open kit →Deflecting customers instead of resolving their problem.
Open kit →A threshold picked once and never revisited — wrong for what a false alarm and a miss actually cost you.
Open kit →Code that passes a benchmark but fails in your codebase.
Open kit →A model metric improving doesn't mean the product got better, and a product metric improving doesn't mean the business benefited. Each layer connects to the next through a hypothesis — and holding those hypotheses is the PM's job. Here is the chain for a RAG assistant; every kit opens with its own.
Tools worth re-running whenever your volume, prices, or pipeline change.
Run it at the start of every new AI request, before anyone scopes a model.
Open →Your draft is saved in this browser — come back and refine it before the review.
Open →Your results are saved in this browser — add them as each version finishes running.
Open →Rerun it every time model prices, volume, or completion rate change.
Open →Rerun it every time someone proposes adding a step.
Open →Run it before every launch review.
Open →Use it when choosing between MAE, RMSE, and MAPE for a forecast.
Open →Plug in this week's numbers; export your picks as .md or .csv.
Open →Four worked cases, each traced from the model metric to the business outcome it missed.
The support bot hit 68% deflection — and repeat contacts rose 9%.
Read the case →ExtractionThe average hid the exact segment deciding whether to renew.
Read the case →CopilotQuality rose from 7.2 to 8.4. Users were editing more, not less.
Read the case →ChatCost per token fell 35%. Lost conversions were worth several times more.
Read the case →Produces: A PRD with context spec, unforgivables, fallbacks, and eval criteria
Open →Produces: Latency, friction, disambiguation, and failure-state designs
Open →Produces: A six-part ask, the "why AI not rules" answer, board framing, response scripts, and an AI SLA
Open →Produces: Business phrasing for every metric
Open →Looking for a specific metric? All 38 metrics with formulas and calculators →