Turbo AI PM
Turbo AI PM

AI Inference Tradeoffs

Understand latency, throughput, and cost well enough to have meaningful technical conversations with engineers. Learn the mental model behind P-01 Latency, P-02 TTFT, and P-03 Throughput — the tradeoffs between them, the levers that move them, and the questions that matter.

Reach for this page when: you're setting an SLO for a new AI feature, someone reports "it feels slow" and you need to know why, or a cost review is coming and engineering wants to change the serving stack.

Compiled from current (2026) LLM inference engineering references. Figures below are industry rough benchmarks, not guarantees — your engineers' numbers on your own hardware and workload are what actually count. Treat specific multipliers as "roughly this shape," not a spec.
01

Two phases, one mental model

Every LLM response happens in two distinct phases. Almost every tradeoff you'll discuss with engineers traces back to one of them — this split is the single highest-leverage thing to internalize.

Prefill
Reading your prompt
What it doesReads the entire prompt in one parallel pass and produces the first token.
BottleneckCompute (GPU FLOPs) — it's doing a lot of math at once.
Drives the metricTTFT — how long until the user sees the first word.
Decode
Writing the answer
What it doesGenerates the rest of the answer one token at a time, looping until done.
BottleneckMemory bandwidth, not compute — each step re-reads model weights and the KV cache. Most production systems are memory-bound here, not FLOPs-bound.
Drives the metricTPOT — how fast tokens stream out after the first one.
Rule of thumb: total wait ≈ TTFT + (TPOT × number of output tokens)

For streaming UIs, TTFT usually matters more — once tokens are visibly appearing, users stay engaged even if the full answer takes a while.

One trap this model doesn't show: in real deployments, TTFT spikes are often not about prefill at all — they're queue wait or a cold GPU spinning up after autoscaling scaled to zero. Before treating a TTFT problem as a prefill-compute problem, ask whether the spikes correlate with autoscaling events rather than prompt length.

02

The metrics that matter

When an engineer says a change "helps throughput but hurts P99 TTFT," this is the vocabulary. Always ask which metric an optimization targets — and what it costs elsewhere.

MetricPlain meaningWhen it's the priority
TTFTDelay before the first word appears. Dominated by prefill plus any queue wait.Interactive chat, voice — anything where a user is staring at a blank screen.
TPOT / ITLAverage gap between streamed tokens during decode. ~6 tokens/sec roughly matches human reading speed.Coding assistants and long responses — tokens must out-pace the reader.
ThroughputTotal tokens the system produces across all concurrent users. Drives cost per token.Batch jobs, agentic workflows, anything serving many users at once.
E2E latencyTotal time to the complete response — TTFT plus all the decode steps.Summarization or jobs that need the full output before the next step.
P95 / P99The slow tail — the 95th/99th percentile request. Averages hide this; the tail is what users complain about.Any SLO. Define targets on percentiles, not means.
03

The central tradeoff: latency vs. throughput

This tension recurs constantly. The lever in the middle is batching — packing multiple requests onto the GPU at the same time.

Bigger batches

Higher throughput

Better GPU utilization, lower cost per token — but can raise TTFT and TPOT, since a new request's heavy prefill can stall the decode steps of requests already in flight.

Smaller batches

Lower latency

Lower latency per user with more dedicated resources — but worse utilization and higher cost per token.

The product question that resolves it: what does the user actually need? Set a latency SLO — e.g. "TTFT under 500ms for 95% of requests, TPOT above 30 tokens/sec" — before engineers tune infrastructure. Without those numbers they can't make good tradeoffs, and you can't tell whether the system is actually working.

A rough way to derive your own target: (what users tolerate before disengaging) × (a safety margin) — measured at p95, not the average. Re-check it after any serving-stack change.
Use caseOptimize forWhy
Real-time chat / support botTTFTA multi-second blank wait kills trust. Happy to serve fewer users well.
Coding assistantTPOT (tolerable TTFT)Devs accept a brief "thinking" pause but need fast streaming after.
Batch summarization / reportsThroughputNobody watches it stream. Higher latency is fine; cost matters more.
Agentic workflow (many chained calls)Throughput + TPOTPer-call latency compounds across chained calls into total task cost.
A short example: the cost win that wasn't
Connecting the dots

A consumer app with an AI chat feature, under pressure to cut infra spend

The change
Infra raised the max batch size to improve GPU utilization. It worked exactly as intended: utilization up, cost per token down 35%. The ticket closed as a clean win.
The system layer
Throughput up. But p95 TTFT went from 600ms to 2.1s — new requests now waited behind the prefill of whatever was already batched. The average TTFT barely moved, which is why the dashboard looked fine.
The product layer
Session abandonment before the first response rose 11%. Messages per session fell. Nobody filed a bug — users just left slightly more often, in the two seconds of blank screen.
The business layer
The feature drove trial-to-paid conversion. The lost conversions were worth several times the infra savings — but the two numbers lived in different dashboards owned by different teams, so for six weeks nobody connected them.
What changed
Batch size was tuned to a p95 TTFT ceiling rather than to utilization alone — landing at roughly 18% cost reduction with TTFT held under 900ms. Abandonment was added as a named guardrail on any serving-config change, and p95 TTFT went on the same dashboard as conversion.

Both teams did their jobs correctly. The gap was that nobody owned the tradeoff between them — which is the PM's job, and the reason latency targets need to exist before infra starts tuning.

04

The optimization toolbox

These are the techniques engineers reach for. They stack — combining several on modern hardware typically delivers a large cost-efficiency improvement over a naive setup. You don't need to implement any of these; you need to know what each one buys and what it costs.

TechniqueLevelEffortQuality risk
System-level — how requests are served
Continuous batchingSystemLowNeutral
KV cache + PagedAttentionSystemLow–MediumNeutral
Speculative decodingSystemMediumEdge case
Disaggregated inferenceSystemHighNeutral
Model-level — changing the model itself
QuantizationModelMediumTradeoff
DistillationModelHighTradeoff
Pruning / sparsityModelHighTradeoff
Application-level — what you send
Prompt cachingApplicationLowNeutral
Context compressionApplicationMediumNeutral
System-level — how requests are served
Continuous (in-flight) batching
Low effortQuality-neutral
Lets new requests join a batch mid-flight instead of waiting for the whole batch to finish. Often the single biggest throughput win — standard in modern serving engines.
Ask: Is our serving engine using continuous batching, or static batches? If static, that's likely the cheapest throughput win available.
KV cache + PagedAttention
Low–Med effortQuality-neutral
The KV cache stores already-computed tokens so they aren't recomputed each step. PagedAttention manages that memory efficiently, cutting waste dramatically. More cache headroom = more concurrent users.
Ask: What's our current cache hit/waste rate, and how much more concurrency could better cache management buy us?
Speculative decoding
Medium effortEdge case
A small, fast "draft" model proposes several tokens; the big model verifies them in one parallel pass. Output is identical to the big model alone, but decode runs meaningfully faster — costs extra memory.
Ask: What's the acceptance rate of the draft model, and what's the memory cost of running two models?
Disaggregated inference
High effortQuality-neutral
Run prefill and decode on separate, specialized hardware — compute-optimized chips for prefill, memory-bandwidth-optimized for decode. An emerging direction showing large throughput gains without a TTFT penalty.
Ask: Is this even available on our current serving stack, or does it require re-architecting?
Model-level — changing the model itself
Quantization
Medium effortQuality tradeoff
Store weights at lower precision (FP16 → FP8/INT8/INT4). Cuts memory significantly and roughly halves cost while keeping most of the accuracy. The main knob with a quality tradeoff to watch.
Ask: What precision, and what's the eval score at that precision vs. FP16 — not just "accuracy is roughly the same"?
Distillation
High effortQuality tradeoff
Train a smaller model to mimic a bigger one. Smaller model = less memory to read per token = lower TPOT and cost, at some quality cost.
Ask: What's the eval delta on our actual eval set, not a generic benchmark, before and after?
Pruning / sparsity
High effortQuality tradeoff
Zero out unimportant weights to speed up the math.
Ask: Has this been validated on our workload, or only on a general benchmark?
Application-level — what you send
Prompt caching
Low effortQuality-neutral
Reuse the prefill work for repeated prefixes (e.g. a long system prompt). Big TTFT and cost win when prompts share boilerplate.
Ask: What share of our prompts share a common prefix, and are we actually caching it?
Context compression
Medium effortQuality-neutral
Don't send tokens the model doesn't need. A meaningful share of input tokens in typical calls is wasted context — and you pay for it twice: once in the bill, again in latency.
Ask: How much of our typical input is boilerplate or unused context we're paying to send every time?
The hidden tax on every speed/cost win: quantization, distillation, and pruning all trade some model quality for speed or cost — that's the "Quality tradeoff" badge above. "Accuracy is roughly the same" is not an eval result — ask to see the actual delta on accuracy, hallucination rate, or whichever quality metric matters for your product, run on your own eval set, before any of these ship. This is the single most common place a speed or cost win turns into a quality surprise three weeks later.
05

The cost framing to carry into any meeting

Inference, not training, is where the money goes

Training happens once; inference runs on every request, forever. The large majority of a deployed model's operational cost is inference.

Per-token prices are falling, but volume is rising faster

Agentic workflows that fire dozens of calls per task turn a cheap per-token price into an expensive per-task cost. Watch cost per task — see C-02 in the field guide.

Optimization gaps exceed hardware-generation gaps

A well-optimized serving stack often beats a naive one by more than a GPU generation does. "Can we tune the serving setup?" is frequently the better question.

06

Questions worth asking your engineers

Seven questions, roughly in the order this page raises them — from diagnosing what's actually slow to checking what a fix costs elsewhere.