TerminologytermStep 6: Quality & riskAllPurchasingRisk

Term: Evaluation (evals)

~6 min read

Estimated time: ~6 min read — for the in-app brief plus opening the primary source.

What this is

Evals are structured tests of whether an AI system meets quality, safety, or business criteria — on tasks and data that matter to you.

Everyday example

Before rolling out an AI contract reviewer, Legal runs 50 real past contracts, scores error types, and only then decides pilot expansion — that is an eval, not a vendor demo.

Evals are tests on real tasks — the only way to know if the tool is good enough for your work.

  • A demo is not an eval. A vendor leaderboard is not your eval.
  • Define pass/fail on your documents, your customers, your risk.
  • Re-run when the model or prompt changes.
  • Include failure cases: hallucination, bias, refusal, latency.

Next action: Before the next rollout, write ten real examples and what “pass” looks like.

What changes in how you lead

How decision rights, process, and ownership should change.

  • No scale without an eval the business owner can explain.
  • Model updates trigger a re-eval, like a change in a control.

Compare related ideas

Evaluation (evals) vs Vendor demo

A demo shows a happy path on the vendor’s sample. An eval uses your tasks, your data distribution, and pre-agreed fail criteria.

Deep dive

Define success in business terms: error rate, time saved, customer impact.

Include edge cases and minority segments where relevant.

Retest after model, prompt, or data changes — ChatGPT / Claude / Grok / Gemini version bumps can move results.

Related terms

Related weekly lessons

terminologyevaluationquality