Services · 15 ways across one gap

All services

Work · proof, not promises

All case studies

Insights · street talk, written down

Magazine

Company · AI in production since 2011

About Voidgap

Evals

evaluations

Systematic tests of an AI system’s outputs against expected results, run automatically after every change.

Evaluationterm 25 of 97guide · 5 stepsAll terms

01In short

Systematic tests of an AI system’s outputs against expected results, run automatically after every change.

In productionNo evals, no go-live. The pass rate is the release criterion.

02Video

03Build your first eval set in five steps

  1. Pick one task and one number.Write down what a good answer means for one process — “claim routed to the right team”, not “the AI is good”. One task, one metric.
  2. Collect real cases.Pull a few hundred real inputs from the process, including the ugly ones: typos, edge cases, the ones people argue about.
  3. Label the expected outcome.Let the people who do the work today mark the right answer. Where they disagree, you found a policy question, not an AI question.
  4. Automate the scoring.Exact match where you can, rules where it fits, LLM-as-a-judge with a written rubric where you must — and spot-check the judge.
  5. Run it on every change.New prompt, new model version, new retrieval setting: the eval runs in CI and the pass rate is the gate to production.

04Checklist

05FAQ

How many cases do we need?

Enough to cover the variety of your process. A few hundred real cases usually beat thousands of synthetic ones.

Can we use a public benchmark instead?

For shortlisting models, yes. To prove the system works in your process, no — benchmarks don’t know your claims, contracts or customers.

Who owns the eval set?

The process owner. Not the vendor, and not the AI team alone — they build it together, the owner signs off the threshold.

From the magazine

Where we help · Secure

Need Evals in production, not on a slide?

Secure