Turbo AI PM
Turbo AI PM

Evaluation Methods

Understand how a real evaluation is built and trusted — not just what a metric means, but what makes the number honest. Classic eval methodology, design pitfalls, and how LLM-as-judge changes (and doesn't change) the picture.

Reach for this page when: you're setting up an eval for the first time, a metric looks suspiciously good, or someone proposes using an LLM to grade your product's own outputs.

This page covers evaluation before you ship. For after — baselines, slices, business fit, and production monitoring — see Production & Business Fit.

Compiled from established ML evaluation practice plus current (2026) LLM-as-judge literature. The classic methodology is stable; the LLM-as-judge section moves faster — treat specific bias findings as "worth checking for," not settled law.
01

What makes an eval trustworthy

Every evaluation rests on three pillars. Each one can quietly lie to you on its own — a great metric run against a bad eval set is still a bad eval.

01

The eval set

What it isThe examples you're testing against.
How it liesStale, too easy, or contaminated by data the model already saw.
02

The metric

What it isHow you turn a response into a number.
How it liesWrong metric for the actual goal, or a number that's easy to game (Goodhart's law — see §5).
03

The judge

What it isWho decides right vs. wrong — a human, a script, or an LLM.
How it liesInconsistent between raters, or systematically biased (see §6).
An eval is only as trustworthy as the weakest of the three pillars — not the average of them.

When a metric looks suspiciously good, check all three before believing it: is the eval set actually hard and uncontaminated, is the metric measuring the thing you care about, and would a second judge agree?

02

Regression metrics

The main field guide covers classification metrics (accuracy, precision, recall, F1). These four are for anything predicting a number, not a class — demand forecasts, price predictions, scoring models.

MetricPlain meaningWhen it's the priority
MAEMean Absolute Error — average absolute difference between predicted and actual, in the original units. Easy to explain to anyone.All errors matter roughly equally, and you want a number stakeholders can sanity-check by eye.
RMSERoot Mean Squared Error — squares errors before averaging, then takes the root. Penalizes large misses much more than small ones.A few big misses are especially costly — you'd rather have many small errors than one huge one.
MAPEMean Absolute Percentage Error — error expressed as a % of the actual value. Scale-independent, so it's comparable across very different-sized targets.Comparing accuracy across targets of very different magnitude — but breaks down when actual values are near zero.
R²Coefficient of determination — how much of the variance in the outcome the model actually explains. 1.0 is perfect; 0 is no better than guessing the average."How much of the story does this model capture," not just "how big is each miss."
See it, don't just read it
Compare the error metrics

Enter a few predicted-vs-actual pairs and watch how MAE, RMSE, and MAPE react to the same errors. Push one prediction far off and see RMSE jump while MAE barely moves — that's the whole reason the metric choice matters.

Predicted
Actual
MAE
—
avg absolute error
RMSE
—
penalizes big misses
MAPE
—
error as % of actual
Your move The metric choice is a business decision disguised as a technical one — and it's yours to frame. "Is one big miss worse than many small ones?" isn't a math question, it's a question about what your users and the business can absorb. A demand forecast that's off by 50% once a quarter (RMSE-costly) is a very different risk than one that's off by 5% every day (MAE-costly). Decide which pain your product can't afford before your team picks the metric to optimize.
03

Choosing between metrics

The precision/recall tradeoff (see Q-02 and Q-03 in the field guide) shows up again one level up, in which curve you trust to summarize a classifier.

Balanced classes

ROC curve / AUC-ROC

Plots true positive rate against false positive rate. Reads cleanly when positives and negatives are roughly evenly represented.

Imbalanced classes

PR curve / PR-AUC

Plots precision against recall. Far more informative when positives are rare (fraud, rare disease, safety flags) — ROC-AUC can look great while precision is actually terrible.

Threshold tuning: most classifiers output a probability, not a yes/no — the threshold you pick to convert one into the other is a product decision, not a technical default. A lower threshold catches more positives (higher recall) at the cost of more false alarms; raise it to protect precision instead. Pick the threshold using the curve, not a single fixed cutoff someone chose once and never revisited.

Your move The threshold is where the metric becomes a user experience. Translate it out of stats and into what the user feels: at this threshold, how many real cases do we miss, and how many false alarms does a user hit before they stop trusting the feature? A fraud flag, a content filter, and a "you might like this" recommender want that dial in completely different places — and that placement is your call to make with the business, not a default to inherit from the model.
04

Eval set design

The eval set is pillar one from §1, and usually the least-questioned part of the whole setup. Four things worth getting right.

01

Build a real golden set

Curated, labeled examples with an agreed-upon correct answer — not just whatever traffic happened to come through. Everything else stands on this.

02

Right-size it

Too small and every score bounces around on noise; too large and iteration slows to a crawl. Size it so a real regression clearly clears the noise floor — not to a round number that felt big enough.

03

Guard against contamination

If the model or judge saw this data during training or tuning, the score is inflated and you won't know by how much. Keep training data and eval data provably separate.

04

Keep a true holdout

Set aside a slice no tuning decision has ever touched, and check it only occasionally. It's the only thing that can catch overfitting to your own eval set (see §5).

Your move You almost certainly won't build the golden set — but you're the one who has to insist it exists, ask who owns it, and push for the hard, ugly, real-user cases to be in it (see survivorship bias in §5). The most useful thing a PM does here is make sure the eval set reflects the messy inputs your actual users send, not the clean ones that were easy to label. If the eval set only contains the easy half of the job, every metric downstream is quietly measuring the wrong thing.
Bootstrapping a golden set from scratch

No users yet, no production logs, nothing to label. This is common — and it's not a dead end. The options in order of effort:

ApproachWhat you produceEffortWatch out for
PM-written examples20–50 hand-crafted input/output pairs, covering the core cases and your expected edge casesHours. The fastest start.Survivorship bias — you'll write the cases you can imagine, not the messy ones real users will send. Seed with deliberately ugly inputs.
Synthetic data generationUse an LLM to generate diverse input variations from a seed set — paraphrases, different tones, edge cases, adversarial inputsHours to a day to set up and review.The generator inherits the biases of the seed. Always human-review a sample. Never use synthetic data as your only eval source.
Expert annotation sprintDomain experts label 100–500 examples; if two+ experts label the same examples, you also get inter-rater agreement as a quality signalDays. The most expensive but highest-quality bootstrap.Experts tend to label what's easy to label. Specifically ask them for edge cases and the hardest examples they can construct.
Shadow mode + early accessRun the model on a limited slice of real traffic before full launch; log outputs for human reviewEngineering setup cost but produces real-distribution dataOnly works if you already have some traffic. Best as a second-phase strategy after a small initial set.
Economics of labeling

Once you have more than a few hundred examples to label, the question of who labels them becomes a cost and quality decision.

Labeling approachCostQualityBest for
Internal team (PMs, domain experts)Expensive in people-hours; "free" in cashHighest — they know the domain and the productAmbiguous edge cases, safety-critical decisions, anything where judgment matters more than speed
Crowdsourcing (e.g. MTurk, Prolific)Low — $0.10–$1 per labelVariable — good for clear tasks, poor for subjective or domain-specific onesLarge-scale labeling of well-defined, objective tasks (is this image safe?, what language is this?)
Managed labeling platforms (e.g. Scale AI, Labelbox)Mid–high — $1–$10+ per labelGood — managed quality control, specialist labelers availableMedium-complexity tasks where you need volume and consistency and can't afford internal expert time
LLM-as-labeler (with human spot-check)Very low — fractions of a cent per labelVaries by task — good at objective tasks, poor at nuanced subjective onesFirst-pass triage, filtering obvious cases before human review, generating label candidates for human approval

In practice: use LLM-assisted labeling for the easy 60%, internal experts for the hard 20%, and a managed platform for the bulk 20% that needs volume but not deep judgment. Never outsource your "unforgivables" (safety, legal, brand) to crowdwork without expert review on top.

05

Common pitfalls

Three failure modes that produce a great-looking number and a worse product.

Goodhart's Law: the metric stops meaning what it used to

You notice firstThe metric's been climbing for weeks, but complaints, churn, or support tickets haven't improved — or got worse. The dashboard and the users disagree.

"When a measure becomes a target, it ceases to be a good measure." Optimizing directly for a metric — rather than the underlying quality it was supposed to stand in for — games the number. Chasing a higher BLEU score, for instance, can inflate lexical overlap with a reference answer without making responses actually better.

Ask the team: what exactly are we optimizing — the proxy, or the real thing it stands for?

Overfitting to your own eval set

You notice firstEval scores keep improving with each iteration, but live/production performance is flat or declining. Great in testing, mediocre in the wild.

Repeatedly tuning against the same fixed eval set eventually fits its specific quirks rather than general quality — the same failure mode as a model overfitting its training data, just one level up, in the human decisions made between iterations. This is exactly what the holdout set in §4 exists to catch.

Ask the team: when did we last check the untouched holdout set — and is it genuinely untouched?

Survivorship bias in what gets tested

You notice firstThe metrics look great, but you keep hearing about real-user failures that "shouldn't be happening" per the numbers — usually on weird, messy, or edge-case inputs.

A golden set built from "clean," well-formed examples hides how the system handles the messy, ambiguous, or malformed inputs that never made it in. If nobody's deliberately adding hard and ugly cases, the eval set quietly drifts toward measuring only the easy part of the job.

Ask the team: what's actually in the eval set — and are the hard, malformed, real-user cases represented?

Your move Goodhart's Law usually starts with a number a PM chose as a target — so you're often the person best placed to catch it. When you set a metric as a goal, you've made it worth gaming; watch for your team hitting the number in ways that don't help users (a hallucination rate that drops because answers got vaguer, a completion rate that rises because the bar for "complete" quietly fell). The healthy habit: pair every target metric with a guardrail metric it shouldn't be allowed to hurt, and keep asking "is the user actually better off, or does the dashboard just look better?"
06

LLM-as-judge

Using a strong LLM to score outputs against a rubric, instead of — or alongside — a human rater. Faster and cheaper than human eval, which is exactly why it needs more guardrails before anyone trusts the number.

First decision: judge, human, or simple metric?

Before any of the mechanics below, the real PM call is whether an LLM judge is even the right tool here. Three options, trading cost and speed against how much you can trust the result:

Simple automated metric

CheapFastCrude
Reach for it whenThere's a clear right answer to match against — exact match, overlap, a number. Near-free and instant, but blind to nuance and meaning.

LLM-as-judge

Mid costScalesNeeds calibration
Reach for it whenQuality is subjective (tone, helpfulness) and you need to score far more examples than humans could. Trustworthy only once calibrated against humans.

Human eval

SlowExpensiveGold standard
Reach for it whenThe stakes are high or the judgment is genuinely hard — launch gates, safety sign-off, anything you'd defend externally. Doesn't scale, so spend it where it counts.

Most mature setups use all three in layers: cheap metrics catch the obvious breaks, an LLM judge covers the subjective middle at scale, and human eval is reserved for the high-stakes calls. It's the same "what does this decision actually need" instinct as setting a latency SLO — match the tool to the stakes, don't default to the fanciest option.

Known biases to test for
BiasWhat happensMitigation
Position biasIn pairwise comparisons ("which is better, A or B"), judges systematically favor whichever answer is shown first — or last — regardless of actual quality.Randomize order, or run both orderings and average the result.
Verbosity biasJudges tend to score longer answers higher, independent of whether the extra length adds anything.Add an explicit length-independence instruction to the rubric, or cap length as a controlled variable.
Self-preference biasA judge model tends to rate outputs from its own model family more favorably.Use a different (ideally stronger) model as judge than the one being evaluated.
Choosing a judge model

Rather than track specific models (which turn over fast), think in two categories: a strong general-purpose frontier model as judge — most capable, but higher cost and latency at scale — or a smaller purpose-built or fine-tuned evaluator model — cheaper and faster, narrower. The tradeoff is capability against cost, and it's the same decision whichever names are current this quarter. Whatever you pick, use a different (ideally stronger) model than the one being evaluated, and see "Judge drift" below for why pinning the version matters more than the specific choice.

Calibrate the judge against humans

A judge that hasn't been checked against human raters is an assumption, not a measurement. Periodically sample judge scores, have humans grade the same outputs blind, and track the agreement rate as its own ongoing metric — not a one-time setup step.

Can I ship on this score? The reliability ladder

A judge score isn't equally trustworthy for every decision. The real PM question isn't "is the judge biased" in the abstract — it's "can I act on this number, or do I need a human in the loop first?" Match the confidence you demand to the stakes of the decision:

Low stakes
Catching regressions in CI. An uncalibrated or roughly-calibrated judge is fine here — you're looking for "did this build get obviously worse," fast, on every commit. A few wrong calls wash out, and the speed is worth more than precision.
Medium stakes
Comparing two approaches, prioritizing work. Trust the judge only once you've checked its human-agreement rate on this kind of task. Report the score with that agreement rate attached, never as bare ground truth.
High stakes
Launch gates, safety or compliance sign-off. Don't ship on a judge score alone. Escalate to human eval for the final call — the judge can pre-filter and narrow what humans review, but it doesn't replace them where the decision is expensive to get wrong.

The move is the same one you make with any metric: translate the number into a decision-confidence level before you act on it, and be honest upward about which rung you're on.

Pairwise vs. pointwise grading
More reliable

Pairwise — "which is better"

Comparing two outputs head-to-head is a much easier, more consistent judgment than rating one output on an absolute scale.

Drifts over time

Pointwise — "rate 1–10"

Feels familiar, but absolute scores drift across sessions and prompts — an 8 today and an 8 next month may not mean the same thing.

Rubric design
Good fit

Fast regression-catching, subjective quality at scale

Catching quality drops in CI between builds, and scoring subjective dimensions (tone, helpfulness) across far more examples than human review could cover.

Not a fit

Final safety sign-off, factual grounding

Not a substitute for grounded factual verification — that's what faithfulness and retrieval checks are for — and not sufficient alone for final high-stakes safety or compliance decisions.

Teaching to the judge: once people know a judge exists, outputs quietly start optimizing for its known biases — padding length, adopting phrasing patterns the judge rewards — rather than actually improving quality. This is Goodhart's Law (§5) applied specifically to LLM judges, and it's the reason judge-human calibration can't be a one-time setup step.
A short example: when a rising judge score meant nothing
Connecting the dots

An AI feature that drafts marketing copy inside a SaaS product

The metric
The team ran an LLM judge scoring draft quality 1–10. After a prompt rewrite, the average score rose from 7.2 to 8.4 — the biggest jump they'd measured. It was reported as a major quality win.
The product layer
Draft acceptance rate: unchanged. Worse, the average number of edits per accepted draft went up, and time-to-publish got slower. Users were doing more work, not less.
What was happening
The rewritten prompt produced noticeably longer drafts. The judge — pointwise, uncalibrated, never checked against human raters — was rewarding length. Classic verbosity bias, straight off the table above. Users had to trim every draft, so the "better" copy cost them more time.
The business layer
This feature was a paid add-on justified by time saved per campaign. The judge said quality was up 17%; the metric the renewal conversation actually rested on had quietly moved the wrong way.
What changed
They switched to pairwise grading against the previous version instead of an absolute score, added a length-independence instruction to the rubric, and started sampling 50 drafts a month for human comparison — making judge-human agreement a tracked metric. The re-scored "improvement" turned out to be roughly flat.

The judge wasn't broken — it was uncalibrated and measuring something nobody had checked. A single question would have caught it: does this score agree with what a human would say?

Judge drift

If the judge is a hosted model the vendor updates silently, your scores can shift for reasons that have nothing to do with your product changing. Pin the judge model version where the provider allows it, and treat a judge-model upgrade as a reason to re-baseline, not a reason to celebrate a sudden score jump.

Your move — reporting it upward When you take a judge score into a launch review or a leadership deck, present it the way you'd present a poll: with its margin of error. A bare "quality is 92%" invites everyone to treat it as ground truth; "92% by our LLM judge, which agrees with human raters ~88% of the time on this task" is the honest version — and it's a move only the PM is positioned to make, because you're the one translating the metric into a decision. Reporting the number without its calibration is how a judge score quietly hardens into a "fact" nobody actually validated.
07

Questions worth asking your eval team

Seven questions, roughly in the order this page raises them.