Turbo AI PM
Kits / workflow

Agents & multi-step pipelines

Tool-using agents and chained AI workflows. Work top to bottom, or jump to the decision in front of you — each one links to the metrics, tool, and playbook that answer it.

The riskEvery step can fail, and failures compound: five steps at 90% complete only 59% of tasks.
Take it to your PRD, Notion, or launch review.

The chain

What has to be true at each layer for this product to create value. The hypotheses between layers are yours to hold — and to check after launch.

If each step is reliable and fast enough, whole tasks complete without a human stepping in.
If tasks complete end to end, people stop doing that work by hand.
BusinessIs the product creating meaningful value?
Hours saved · cost per completed task vs. human cost

↔ marks each metric's shadow metric — the one to watch alongside it so the number can't improve while the product gets worse.

01

Before you build

Should this be AI at all — and which capability fits?

Before scoping a model, check that the problem has a pattern to learn, data you can reach at decision time, and a cost of being wrong you can live with. If a rule would get most of the value, the rule is the better product.

For Agents & multi-step pipelinesThe test for an agent: is the process already written down and stable? Automating an undocumented process usually means automating a mess — and each step multiplies the failure rate.
Evidence, business implication, deeper reading
Evidence you need: Examples of the decision being made today, where the data lives and whether it is available at decision time, and an honest estimate of what a wrong answer costs.
Business implication: Choosing AI for a problem rules could solve buys you cost, latency, and a maintenance burden with no extra value.

Check whether the unit economics can work

Estimate cost per completed task against the value of that task before building — margins are set by architecture, not tuning.

Evidence, business implication, deeper reading
Evidence you need: Volume estimate, token estimate per task, model pricing, and the human cost (or revenue) per task.
Business implication: A feature that works but destroys gross margin is a failed product.

Make the case to build it

Ask for one specific decision by a specific date: the problem in their units, why AI and not rules, the expected value with its assumptions visible, the risk said out loud, and the checkpoint you will report back at.

Evidence, business implication, deeper reading
Evidence you need: The baseline (what today costs and achieves), a value range with the two assumptions that move it most, cost per completed task vs. the human cost, the rules-vs-model comparison, and the top failure modes with their mitigations.
Business implication: A named checkpoint turns a bet into a staged decision — much easier to approve than an open-ended one. For this product: Hours saved · cost per completed task vs. human cost.

Define what the model may do, must never do, and needs to know

Write the operating envelope before anyone writes a prompt: context the model sees, unforgivables, and what happens when it fails.

For Agents & multi-step pipelinesList every tool the agent can call and which actions require confirmation.
Evidence, business implication, deeper reading
Evidence you need: Agreed context spec, a written list of unforgivables, and a fallback for each failure mode.
Business implication: Scope decided here drives both cost and risk for the life of the feature.

Decide how many steps the pipeline can afford

End-to-end success is the product of every step. Fewer, more reliable steps usually beat more capable, longer chains.

Evidence, business implication, deeper reading
Evidence you need: Estimated reliability per step and cost per call.
Business implication: Every failed task still costs tokens — cost per successful task is the real unit cost.
02

Evaluate

Evidence, business implication, deeper reading
Evidence you need: The starter metrics below, each paired with the metric that catches it being gamed.
Business implication: The layer-3 metric is the one leadership will ask about — name it now. For this product: Hours saved · cost per completed task vs. human cost.

Establish a baseline worth beating

An agent has to beat simpler versions of itself, not just the manual process. Define what a completed task is, then compare four versions on the same tasks.

For Agents & multi-step pipelines
  1. Define success first. What counts as a completed task — and does a human stepping in mid-task count as a failure? Decide before measuring anything.
  2. Baseline each step. A simple rule against that step's model, to find the weakest link before you chain anything.
  3. Compare four versions on one task set: the manual process (timed), a single well-written prompt, a 2–3 step version, and the full pipeline.
Evidence, business implication, deeper reading
Evidence you need: Your written definition of a completed task, a rule-vs-model check on each step, then completion rate and cost per completed task for all four versions — the manual process, a single prompt, a 2–3 step version, and the full pipeline — on the same task set, with enough tasks to tell them apart.
Business implication: Every step you add multiplies the chance of failure and adds cost — and failed tasks still spend tokens. If a single prompt gets most of the value, the extra steps are cost with no return.

Design an evaluation you can trust

Representative data, a scoring method that matches the task, and a holdout nobody tunes against.

Evidence, business implication, deeper reading
Evidence you need: A golden set with an owner and a size target, slices defined up front, and a judge calibrated against humans.
Business implication: A weak eval means you ship blind and cannot explain failures afterward.
03

Ship

Decide whether it is ready to ship

The bar depends on blast radius and reversibility, not on the accuracy number alone.

For Agents & multi-step pipelinesJudge the end-to-end rate, not step accuracy. Any irreversible action (payments, deletes, sends) moves you to the red quadrant.
Evidence, business implication, deeper reading
Evidence you need: Your quadrant, the threshold for it, the eval result against the baseline, a working fallback, and the numbers you promised at the funding checkpoint.
Business implication: A documented bar turns post-launch blame into post-launch learning.

Design what users see when it is wrong

Every AI flow needs a designed failure state: specific copy, a way to recover, and a human path.

For Agents & multi-step pipelinesConfirm before irreversible steps; on failure, hand off with the work so far preserved.
Evidence, business implication, deeper reading
Evidence you need: Fallback copy for timeouts, low confidence, out-of-scope requests, and refusals.
Business implication: Users forgive visible, recoverable failures; they abandon silent ones.
04

In production

Set up monitoring that separates "the model changed" from "the world changed"

Alert on breakage, review drift daily, review quality weekly, review business impact monthly.

Evidence, business implication, deeper reading
Evidence you need: Error and latency alerts, input-mix and confidence drift, slice quality, and the layer-3 metric on a monthly cadence.
Business implication: Degradation you catch in a week costs far less than one you find in a quarterly review.

Turn production signal into a better next version

Design what each interaction captures and where it goes — implicit, explicit, and outcome signal feeding the eval set.

For Agents & multi-step pipelinesCapture where tasks fail, by step, so you fix the weakest link first.
Evidence, business implication, deeper reading
Evidence you need: A named owner for the signal pipeline and a path from user action to golden set.
Business implication: The flywheel is the only part of an AI feature that compounds over time.
!

Something looks wrong

Start from what you're seeing. Each symptom points to its most likely cause and the metrics that confirm it.

Every step tests well, but tasks fail
Compounding: step reliabilities multiply.
Cost per task spikes
Retries, loops, or runaway tool calls on hard inputs.
The agent did something it should not have
A missing confirmation step or an unwritten unforgivable.

Other kits