Term: Evaluation (evals)
~6 min read
Estimated time: ~6 min read — for the in-app brief plus opening the primary source.
What this is
Evals are structured tests of whether an AI system meets quality, safety, or business criteria — on tasks and data that matter to you.
Everyday example
Before rolling out an AI contract reviewer, Legal runs 50 real past contracts, scores error types, and only then decides pilot expansion — that is an eval, not a vendor demo.
Evals are tests on real tasks — the only way to know if the tool is good enough for your work.
- A demo is not an eval. A vendor leaderboard is not your eval.
- Define pass/fail on your documents, your customers, your risk.
- Re-run when the model or prompt changes.
- Include failure cases: hallucination, bias, refusal, latency.
Next action: Before the next rollout, write ten real examples and what “pass” looks like.
What changes in how you lead
How decision rights, process, and ownership should change.
- No scale without an eval the business owner can explain.
- Model updates trigger a re-eval, like a change in a control.
Compare related ideas
Evaluation (evals) vs Vendor demo
A demo shows a happy path on the vendor’s sample. An eval uses your tasks, your data distribution, and pre-agreed fail criteria.
Deep dive
Define success in business terms: error rate, time saved, customer impact.
Include edge cases and minority segments where relevant.
Retest after model, prompt, or data changes — ChatGPT / Claude / Grok / Gemini version bumps can move results.