What makes an eval trustworthy
Every evaluation rests on three pillars. Each one can quietly lie to you on its own — a great metric run against a bad eval set is still a bad eval.
The eval set
The metric
The judge
When a metric looks suspiciously good, check all three before believing it: is the eval set actually hard and uncontaminated, is the metric measuring the thing you care about, and would a second judge agree?
Regression metrics
The main field guide covers classification metrics (accuracy, precision, recall, F1). These four are for anything predicting a number, not a class — demand forecasts, price predictions, scoring models.
| Metric | Plain meaning | When it's the priority |
|---|---|---|
| MAE | Mean Absolute Error — average absolute difference between predicted and actual, in the original units. Easy to explain to anyone. | All errors matter roughly equally, and you want a number stakeholders can sanity-check by eye. |
| RMSE | Root Mean Squared Error — squares errors before averaging, then takes the root. Penalizes large misses much more than small ones. | A few big misses are especially costly — you'd rather have many small errors than one huge one. |
| MAPE | Mean Absolute Percentage Error — error expressed as a % of the actual value. Scale-independent, so it's comparable across very different-sized targets. | Comparing accuracy across targets of very different magnitude — but breaks down when actual values are near zero. |
| R² | Coefficient of determination — how much of the variance in the outcome the model actually explains. 1.0 is perfect; 0 is no better than guessing the average. | "How much of the story does this model capture," not just "how big is each miss." |
Choosing between metrics
The precision/recall tradeoff (see Q-02 and Q-03 in the field guide) shows up again one level up, in which curve you trust to summarize a classifier.
ROC curve / AUC-ROC
Plots true positive rate against false positive rate. Reads cleanly when positives and negatives are roughly evenly represented.
PR curve / PR-AUC
Plots precision against recall. Far more informative when positives are rare (fraud, rare disease, safety flags) — ROC-AUC can look great while precision is actually terrible.
Threshold tuning: most classifiers output a probability, not a yes/no — the threshold you pick to convert one into the other is a product decision, not a technical default. A lower threshold catches more positives (higher recall) at the cost of more false alarms; raise it to protect precision instead. Pick the threshold using the curve, not a single fixed cutoff someone chose once and never revisited.
Eval set design
The eval set is pillar one from §1, and usually the least-questioned part of the whole setup. Four things worth getting right.
Build a real golden set
Curated, labeled examples with an agreed-upon correct answer — not just whatever traffic happened to come through. Everything else stands on this.
Right-size it
Too small and every score bounces around on noise; too large and iteration slows to a crawl. Size it so a real regression clearly clears the noise floor — not to a round number that felt big enough.
Guard against contamination
If the model or judge saw this data during training or tuning, the score is inflated and you won't know by how much. Keep training data and eval data provably separate.
Keep a true holdout
Set aside a slice no tuning decision has ever touched, and check it only occasionally. It's the only thing that can catch overfitting to your own eval set (see §5).
No users yet, no production logs, nothing to label. This is common — and it's not a dead end. The options in order of effort:
| Approach | What you produce | Effort | Watch out for |
|---|---|---|---|
| PM-written examples | 20–50 hand-crafted input/output pairs, covering the core cases and your expected edge cases | Hours. The fastest start. | Survivorship bias — you'll write the cases you can imagine, not the messy ones real users will send. Seed with deliberately ugly inputs. |
| Synthetic data generation | Use an LLM to generate diverse input variations from a seed set — paraphrases, different tones, edge cases, adversarial inputs | Hours to a day to set up and review. | The generator inherits the biases of the seed. Always human-review a sample. Never use synthetic data as your only eval source. |
| Expert annotation sprint | Domain experts label 100–500 examples; if two+ experts label the same examples, you also get inter-rater agreement as a quality signal | Days. The most expensive but highest-quality bootstrap. | Experts tend to label what's easy to label. Specifically ask them for edge cases and the hardest examples they can construct. |
| Shadow mode + early access | Run the model on a limited slice of real traffic before full launch; log outputs for human review | Engineering setup cost but produces real-distribution data | Only works if you already have some traffic. Best as a second-phase strategy after a small initial set. |
Once you have more than a few hundred examples to label, the question of who labels them becomes a cost and quality decision.
| Labeling approach | Cost | Quality | Best for |
|---|---|---|---|
| Internal team (PMs, domain experts) | Expensive in people-hours; "free" in cash | Highest — they know the domain and the product | Ambiguous edge cases, safety-critical decisions, anything where judgment matters more than speed |
| Crowdsourcing (e.g. MTurk, Prolific) | Low — $0.10–$1 per label | Variable — good for clear tasks, poor for subjective or domain-specific ones | Large-scale labeling of well-defined, objective tasks (is this image safe?, what language is this?) |
| Managed labeling platforms (e.g. Scale AI, Labelbox) | Mid–high — $1–$10+ per label | Good — managed quality control, specialist labelers available | Medium-complexity tasks where you need volume and consistency and can't afford internal expert time |
| LLM-as-labeler (with human spot-check) | Very low — fractions of a cent per label | Varies by task — good at objective tasks, poor at nuanced subjective ones | First-pass triage, filtering obvious cases before human review, generating label candidates for human approval |
In practice: use LLM-assisted labeling for the easy 60%, internal experts for the hard 20%, and a managed platform for the bulk 20% that needs volume but not deep judgment. Never outsource your "unforgivables" (safety, legal, brand) to crowdwork without expert review on top.
Common pitfalls
Three failure modes that produce a great-looking number and a worse product.
Goodhart's Law: the metric stops meaning what it used to
"When a measure becomes a target, it ceases to be a good measure." Optimizing directly for a metric — rather than the underlying quality it was supposed to stand in for — games the number. Chasing a higher BLEU score, for instance, can inflate lexical overlap with a reference answer without making responses actually better.
Ask the team: what exactly are we optimizing — the proxy, or the real thing it stands for?
Overfitting to your own eval set
Repeatedly tuning against the same fixed eval set eventually fits its specific quirks rather than general quality — the same failure mode as a model overfitting its training data, just one level up, in the human decisions made between iterations. This is exactly what the holdout set in §4 exists to catch.
Ask the team: when did we last check the untouched holdout set — and is it genuinely untouched?
Survivorship bias in what gets tested
A golden set built from "clean," well-formed examples hides how the system handles the messy, ambiguous, or malformed inputs that never made it in. If nobody's deliberately adding hard and ugly cases, the eval set quietly drifts toward measuring only the easy part of the job.
Ask the team: what's actually in the eval set — and are the hard, malformed, real-user cases represented?
LLM-as-judge
Using a strong LLM to score outputs against a rubric, instead of — or alongside — a human rater. Faster and cheaper than human eval, which is exactly why it needs more guardrails before anyone trusts the number.
Before any of the mechanics below, the real PM call is whether an LLM judge is even the right tool here. Three options, trading cost and speed against how much you can trust the result:
Simple automated metric
LLM-as-judge
Human eval
Most mature setups use all three in layers: cheap metrics catch the obvious breaks, an LLM judge covers the subjective middle at scale, and human eval is reserved for the high-stakes calls. It's the same "what does this decision actually need" instinct as setting a latency SLO — match the tool to the stakes, don't default to the fanciest option.
| Bias | What happens | Mitigation |
|---|---|---|
| Position bias | In pairwise comparisons ("which is better, A or B"), judges systematically favor whichever answer is shown first — or last — regardless of actual quality. | Randomize order, or run both orderings and average the result. |
| Verbosity bias | Judges tend to score longer answers higher, independent of whether the extra length adds anything. | Add an explicit length-independence instruction to the rubric, or cap length as a controlled variable. |
| Self-preference bias | A judge model tends to rate outputs from its own model family more favorably. | Use a different (ideally stronger) model as judge than the one being evaluated. |
Rather than track specific models (which turn over fast), think in two categories: a strong general-purpose frontier model as judge — most capable, but higher cost and latency at scale — or a smaller purpose-built or fine-tuned evaluator model — cheaper and faster, narrower. The tradeoff is capability against cost, and it's the same decision whichever names are current this quarter. Whatever you pick, use a different (ideally stronger) model than the one being evaluated, and see "Judge drift" below for why pinning the version matters more than the specific choice.
A judge that hasn't been checked against human raters is an assumption, not a measurement. Periodically sample judge scores, have humans grade the same outputs blind, and track the agreement rate as its own ongoing metric — not a one-time setup step.
A judge score isn't equally trustworthy for every decision. The real PM question isn't "is the judge biased" in the abstract — it's "can I act on this number, or do I need a human in the loop first?" Match the confidence you demand to the stakes of the decision:
The move is the same one you make with any metric: translate the number into a decision-confidence level before you act on it, and be honest upward about which rung you're on.
Pairwise — "which is better"
Comparing two outputs head-to-head is a much easier, more consistent judgment than rating one output on an absolute scale.
Pointwise — "rate 1–10"
Feels familiar, but absolute scores drift across sessions and prompts — an 8 today and an 8 next month may not mean the same thing.
- Be specific, not vague. "Is this a good answer?" produces inconsistent, hard-to-audit scores.
- Decompose the dimensions. Score faithfulness, tone, and completeness separately — a single blended score hides which one actually moved.
- Use a different or stronger judge model than the one under evaluation, specifically to avoid self-preference bias.
- Budget for cost and latency — judge calls at scale are still real inference spend (see the cost framing in Inference Tradeoffs).
Fast regression-catching, subjective quality at scale
Catching quality drops in CI between builds, and scoring subjective dimensions (tone, helpfulness) across far more examples than human review could cover.
Final safety sign-off, factual grounding
Not a substitute for grounded factual verification — that's what faithfulness and retrieval checks are for — and not sufficient alone for final high-stakes safety or compliance decisions.
An AI feature that drafts marketing copy inside a SaaS product
The judge wasn't broken — it was uncalibrated and measuring something nobody had checked. A single question would have caught it: does this score agree with what a human would say?
If the judge is a hosted model the vendor updates silently, your scores can shift for reasons that have nothing to do with your product changing. Pin the judge model version where the provider allows it, and treat a judge-model upgrade as a reason to re-baseline, not a reason to celebrate a sudden score jump.
Questions worth asking your eval team
Seven questions, roughly in the order this page raises them.
- →What's actually in our eval set, and when was it last refreshed?
- →Could our model or judge have seen this eval data during training or tuning?
- →Are we tracking judge-human agreement, or just trusting the judge's number?
- →Is this eval pairwise or pointwise — and do we know why we picked that?
- →What's our rubric actually measuring, dimension by dimension?
- →Have we checked for position, verbosity, or self-preference bias in the judge's scores?
- →Is anyone optimizing directly against this metric in a way that could be gaming it?