Two phases, one mental model
Every LLM response happens in two distinct phases. Almost every tradeoff you'll discuss with engineers traces back to one of them — this split is the single highest-leverage thing to internalize.
For streaming UIs, TTFT usually matters more — once tokens are visibly appearing, users stay engaged even if the full answer takes a while.
One trap this model doesn't show: in real deployments, TTFT spikes are often not about prefill at all — they're queue wait or a cold GPU spinning up after autoscaling scaled to zero. Before treating a TTFT problem as a prefill-compute problem, ask whether the spikes correlate with autoscaling events rather than prompt length.
The metrics that matter
When an engineer says a change "helps throughput but hurts P99 TTFT," this is the vocabulary. Always ask which metric an optimization targets — and what it costs elsewhere.
| Metric | Plain meaning | When it's the priority |
|---|---|---|
| TTFT | Delay before the first word appears. Dominated by prefill plus any queue wait. | Interactive chat, voice — anything where a user is staring at a blank screen. |
| TPOT / ITL | Average gap between streamed tokens during decode. ~6 tokens/sec roughly matches human reading speed. | Coding assistants and long responses — tokens must out-pace the reader. |
| Throughput | Total tokens the system produces across all concurrent users. Drives cost per token. | Batch jobs, agentic workflows, anything serving many users at once. |
| E2E latency | Total time to the complete response — TTFT plus all the decode steps. | Summarization or jobs that need the full output before the next step. |
| P95 / P99 | The slow tail — the 95th/99th percentile request. Averages hide this; the tail is what users complain about. | Any SLO. Define targets on percentiles, not means. |
The central tradeoff: latency vs. throughput
This tension recurs constantly. The lever in the middle is batching — packing multiple requests onto the GPU at the same time.
Higher throughput
Better GPU utilization, lower cost per token — but can raise TTFT and TPOT, since a new request's heavy prefill can stall the decode steps of requests already in flight.
Lower latency
Lower latency per user with more dedicated resources — but worse utilization and higher cost per token.
The product question that resolves it: what does the user actually need? Set a latency SLO — e.g. "TTFT under 500ms for 95% of requests, TPOT above 30 tokens/sec" — before engineers tune infrastructure. Without those numbers they can't make good tradeoffs, and you can't tell whether the system is actually working.
| Use case | Optimize for | Why |
|---|---|---|
| Real-time chat / support bot | TTFT | A multi-second blank wait kills trust. Happy to serve fewer users well. |
| Coding assistant | TPOT (tolerable TTFT) | Devs accept a brief "thinking" pause but need fast streaming after. |
| Batch summarization / reports | Throughput | Nobody watches it stream. Higher latency is fine; cost matters more. |
| Agentic workflow (many chained calls) | Throughput + TPOT | Per-call latency compounds across chained calls into total task cost. |
A consumer app with an AI chat feature, under pressure to cut infra spend
Both teams did their jobs correctly. The gap was that nobody owned the tradeoff between them — which is the PM's job, and the reason latency targets need to exist before infra starts tuning.
The optimization toolbox
These are the techniques engineers reach for. They stack — combining several on modern hardware typically delivers a large cost-efficiency improvement over a naive setup. You don't need to implement any of these; you need to know what each one buys and what it costs.
| Technique | Level | Effort | Quality risk |
|---|---|---|---|
| System-level — how requests are served | |||
| Continuous batching | System | Low | Neutral |
| KV cache + PagedAttention | System | Low–Medium | Neutral |
| Speculative decoding | System | Medium | Edge case |
| Disaggregated inference | System | High | Neutral |
| Model-level — changing the model itself | |||
| Quantization | Model | Medium | Tradeoff |
| Distillation | Model | High | Tradeoff |
| Pruning / sparsity | Model | High | Tradeoff |
| Application-level — what you send | |||
| Prompt caching | Application | Low | Neutral |
| Context compression | Application | Medium | Neutral |
The cost framing to carry into any meeting
Inference, not training, is where the money goes
Training happens once; inference runs on every request, forever. The large majority of a deployed model's operational cost is inference.
Per-token prices are falling, but volume is rising faster
Agentic workflows that fire dozens of calls per task turn a cheap per-token price into an expensive per-task cost. Watch cost per task — see C-02 in the field guide.
Optimization gaps exceed hardware-generation gaps
A well-optimized serving stack often beats a naive one by more than a GPU generation does. "Can we tune the serving setup?" is frequently the better question.
Questions worth asking your engineers
Seven questions, roughly in the order this page raises them — from diagnosing what's actually slow to checking what a fix costs elsewhere.
- →Is this workload prefill-heavy or decode-heavy? (Long prompts vs. long outputs change everything.)
- →Are our slow requests actually prefill-bound, or is it queue wait / a cold start from autoscaling?
- →What's our TTFT and TPOT target, and at what percentile?
- →Are we memory-bound or compute-bound at decode right now?
- →What does this optimization cost on the other metrics?
- →What's the eval delta on our own data before this ships — not just "accuracy is roughly the same"?
- →What's our cost per task, not just per token, once we count retries and agent loops?