Evals
The tests that say whether the thing actually works for your job.
An eval is a test set with a scoring method: a body of representative tasks, the answers you would accept, and a way of marking. Public leaderboards are evals on generic tasks — useful for tracking the frontier, close to useless for deciding whether a system will handle your correspondence, your legislation, or your complaints.
The evals that matter are the ones built from your own material: a hundred real cases, the answers a competent officer would give, run again every time the model, prompt or vendor changes. This is unglamorous and it is the single highest-leverage thing a technical team can build, because without it every upgrade is a guess and every vendor claim is unfalsifiable.
Building one is less work than it sounds. A hundred cases with expected outputs is a week for someone who knows the domain, and it pays for itself the first time a vendor pushes a model update.
Why it matters here
This is the difference between a pilot that ends in a decision and a pilot that ends in a feeling. It is also the artefact that makes a procurement defensible: you can say what you tested, on what, and what the result was.
The question to ask
Show me your eval set. If there is not one, what are we relying on instead?
Reviewed 2026-09-20 · All decoders