Evals Interview Questions (2026)
Covers AI evals — measuring whether a system actually works, on your task. See also all interview topics. These assume you already know the concept — if this is unfamiliar, read the full concept page first; the questions test judgment on top of the concept, not the concept itself.
AI Evals
A new model version scores higher than your current one on every public benchmark you can find. Is that enough to justify upgrading in production?
No. Public benchmarks measure broad capability, not how the model behaves on your specific task and data. The only way to know is running it against your own eval set and comparing real outputs on real examples from your domain — a model that's better on average can still be worse for you.
Your eval pipeline uses an LLM to score whether each response is "good." The judge model gives noticeably higher scores to longer answers, even when a shorter one was actually more correct. What's happening?
A known bias in LLM-as-judge scoring — judges tend to favor verbosity over precision. This is exactly why judge scores need periodic calibration against real human review rather than being trusted on their own; a judge model isn't a neutral oracle, it has its own systematic tendencies.
You're deciding how to test a support-ticket classifier. Should you rely entirely on an LLM judge, entirely on exact-match checks, or something else?
Layer them. Deterministic checks — does the output match one of the five valid categories, is the JSON well-formed — catch cheap, mechanical failures for free. LLM-as-judge handles what exact-match can't: whether a free-text response is actually reasonable. Human review calibrates the judge periodically, since neither layer alone covers everything.
Your agent completes a multi-step task and the final answer looks correct. Is checking the final output enough to trust the eval?
Not necessarily. An agent that reaches the right answer through an unreliable or unsafe path — calling the wrong tool first, skipping a check it should have made — can get lucky on this run and fail differently on the next one. Evaluating an agent often means scoring the path it took, not only whether the last thing it said was right.