Services · 15 ways across one gap

All services

Work · proof, not promises

All case studies

Insights · street talk, written down

Magazine

Company · AI in production since 2011

About Voidgap
Sample data Every system and every number on this board is fictional, except Ask Voidgap, whose eval card runs live on /ask. Real client data only with written consent.

Open evals

Pass rate · gate · cost · incidents

Sample board · as of September 2026

Evals, or it
didn’t happen.

A public board for AI in production: how often it passes its test set, the bar it has to clear, what a task costs, and when it last went wrong.

Systems on the board5
Above the gate3
Below the gate1
Real, not a sample1

sample Four fictional systems and one real one, Ask Voidgap.

01 · The board

Five systems,
nothing hidden.

Hover or tab into a chart to read each run. The dashed line is the gate: below it, nothing ships.

Runs
Status

The real one · on this site

Ask Voidgap

in place

Answers questions from the pages of this website and cites them. Its eval card is public: eight questions, the page each answer must cite, one it must refuse. It runs on your account, in your browser, when you press the button.

Last incident

incident log · to decide

Last run
–/8yours, when you run it
Gate
to decide
Cost per task
to decide

No run history yet. Runs are not stored anywhere, so there is no line to draw. A public history needs storage and a decision: run history · to decide

Run the eval card

Drafts for a human agent

Support reply drafts

sample
Last run
93.1 %above the gate
Gate
90 %380 test cases
Cost per task
€0.02per draft

Pass means Facts correct, right tone, no promise outside the policy

80859095100gate 90 %93.16 Jul10 Aug21 Sep

Last incident

Drafts quoted an outdated returns window after a policy change. Knowledge base updated; three cases added to the test set.

Data table
RunPass ratevs gateNote
21 Sep 202693.1 %+3.1
14 Sep 202692.9 %+2.9
7 Sep 202693.2 %+3.2
31 Aug 202693.6 %+3.6
24 Aug 202693.0 %+3.0
17 Aug 202693.3 %+3.3
10 Aug 202692.8 %+2.8
3 Aug 202686.5 %−3.5Model version update scored below the gate. Not deployed.
27 Jul 202694.0 %+4.0
20 Jul 202693.5 %+3.5
13 Jul 202693.9 %+3.9
6 Jul 202693.2 %+3.2

Documents into the ERP

Delivery note extraction

sample
Last run
97.6 %above the gate
Gate
97 %520 test cases
Cost per task
€0.03per document

Pass means Every required field correct

9095100gate 97 %97.66 Jul10 Aug21 Sep

Last incident

A scanner update rotated pages. Confidence dropped, 212 documents went to manual review, none reached the ERP unchecked.

Data table
RunPass ratevs gateNote
21 Sep 202697.6 %+0.6
14 Sep 202697.5 %+0.5
7 Sep 202697.1 %+0.1
31 Aug 202697.4 %+0.4
24 Aug 202697.2 %+0.2First run above the gate. Went live the week after.
17 Aug 202696.9 %−0.1
10 Aug 202696.7 %−0.3
3 Aug 202696.1 %−0.9
27 Jul 202696.3 %−0.7
20 Jul 202696.0 %−1.0
13 Jul 202695.4 %−1.6
6 Jul 202695.6 %−1.4

Flags risky clauses for legal

Contract clause finder

sample
Last run
95.2 %below the gate
Gate
98 %150 test cases
Cost per task
€0.31per contract

Pass means Every clause on the checklist found, nothing invented

859095100gate 98 %95.26 Jul10 Aug21 Sep

Last incident

None in production. It is not in production: below the gate it runs in shadow mode, next to the lawyers, not instead of them.

Data table
RunPass ratevs gateNote
21 Sep 202695.2 %−2.8
14 Sep 202695.0 %−3.0
7 Sep 202694.8 %−3.2
31 Aug 202694.6 %−3.4
24 Aug 202694.1 %−3.9
17 Aug 202693.9 %−4.1
10 Aug 202693.5 %−4.5
3 Aug 202693.0 %−5.0
27 Jul 202692.4 %−5.6
20 Jul 202692.6 %−5.4
13 Jul 202692.1 %−5.9
6 Jul 202691.8 %−6.2

Voice agent

Appointment booking by phone

sample
Last run
94.4 %above the gate
Gate
92 %260 test cases
Cost per task
€0.18per call

Pass means Right slot, right person, caller told it is an AI

9095100gate 92 %94.46 Jul10 Aug21 Sep

Last incident

Callers with double-barrelled surnames were booked under half their name. 23 bookings corrected by hand; 40 name cases added to the test set.

Data table
RunPass ratevs gateNote
21 Sep 202694.4 %+2.4
14 Sep 202694.8 %+2.8
7 Sep 202694.6 %+2.6
31 Aug 202694.3 %+2.3
24 Aug 202694.7 %+2.7
17 Aug 202694.9 %+2.9
10 Aug 202694.2 %+2.2
3 Aug 202694.6 %+2.6
27 Jul 202694.8 %+2.8
20 Jul 202694.3 %+2.3
13 Jul 202694.5 %+2.5
6 Jul 202694.0 %+2.0

Weekly runs on a fixed test set. A run below the gate blocks the deploy; the line shows what was tested, not what went live. Hollow dots mark runs with a note.

02 · The four numbers

Four numbers
per system.

If a vendor cannot give you these four, the system is a demo with a login.

  1. 1Pass rate

    Share of the test set that passes on a run. The test set comes from real cases, with the expected result written down before the build.

  2. 2The gate

    The pass rate the owner agreed to before launch. Below it, a change does not deploy. It is set per system, because 90 % is fine for a draft and not for a customs value.

  3. 3Cost per task

    Model and infrastructure cost per finished task, not per token. The number finance can put next to the salary it saves or the hours it frees.

  4. 4Last incident

    The last time it went wrong for a real user, with the date and what changed after. A board without incidents is either very new or not honest.

When the gate is missing: Ship happens

Before the build

Agree the
gate first.

Every project starts with the test set and the number it has to beat. Then we build, and you can watch the line.