AI Evaluation

How to tell whether an AI system is actually working, rather than just looking like it is, and how to keep telling once anything about it changes.

  • AI Evals — an eval is a repeatable test that measures how an AI system performs on your task, using real examples and a scoring method — not a gut feeling about a few outputs.
  • LLM-as-Judge — uses a model to score another model's output against a rubric, making it possible to evaluate subjective quality at scale.