AI Evals

A model can score near the top of every public benchmark and still get a noticeable share of answers wrong on your actual task, with your actual data. Benchmarks measure broad capability. They don't measure whether this model, on your prompts, doing your job, is any good — that's what an eval is for.

An eval is a repeatable test: a fixed set of real example inputs, run through your system, scored by a defined method. Change a prompt, switch models, or adjust a pipeline, and you run the same test again and compare the score — instead of trying both versions a few times and going with whichever felt better.

Why "it seemed fine" isn't enough

Trying a handful of prompts and judging the output by eye doesn't scale, doesn't catch a regression a change introduced somewhere you weren't looking, and doesn't survive a model upgrade — what felt fine on the five examples you tried can quietly get worse on the hundred you didn't. An eval turns "does this work" from an impression into a number you can track over time and compare across changes.

Three layers, used together

Deterministic checks. Does the output match one of the allowed categories? Is the JSON well-formed? These are cheap, exact, and catch mechanical failures for free — but they only work when there's a single correct answer to check against.

LLM-as-judge. For everything else — is this free-text answer actually reasonable, does it faithfully reflect the source, is the tone right — a second model call scores the output against a rubric. This is what makes evaluating subjective quality possible at scale, since a person can't manually read thousands of outputs every time something changes.

Human review. Used to spot-check and calibrate the judge, not to review everything. A judge model isn't a neutral oracle — it has its own tendencies, including favoring longer, more verbose answers over shorter, more precise ones, and it can drift over time. Periodic human review is what catches the judge itself going wrong.

Comparing prompts, model versions, and agents

Comparing two prompt versions. Run both against the same 50 real support tickets and score each, instead of trying each once and picking whichever answer reads better in the moment.

Deciding whether to upgrade a model. A newer model scoring higher on public benchmarks doesn't tell you how it behaves on your domain — running your own eval set against both versions does.

Scoring an AI agent, not just its final answer. An agent that reaches the right answer through an unreliable path — calling the wrong tool first, skipping a check it should have made — can get lucky this run and fail differently the next. Evaluating an agent often means scoring the steps it took, not only the last thing it said.

When you don't need this

A one-off task, or something a person is already checking by hand every time it runs, doesn't need an eval pipeline built around it — that's real infrastructure, worth the cost once a prompt or pipeline is being changed repeatedly and shipped to real users, not before.

In this guide
  1. Why "it seemed fine" isn't enough
  2. Three layers, used together
  3. Comparing prompts, model versions, and agents
  4. When you don't need this
  5. FAQ

FAQ

If a new model version scores higher on every public benchmark, do you still need to run your own evals before upgrading?

Yes. Public benchmarks measure broad capability across generic tasks, not how a model behaves on your specific prompts and data. A model that's genuinely better on average can still perform worse on your particular task — the only way to know is running both versions against your own eval set.

Can an LLM judge fully replace a human reviewer?

No. Judge models have documented biases — favoring longer, more verbose answers over shorter correct ones, sensitivity to the order options are presented in, and drift in scoring behavior over time. Treat a judge as a scalable first pass, calibrated periodically against real human review, not a neutral final word.

Practice interview questions on AI evals →