Sample data Every system and every number on this board is fictional, except Ask Voidgap, whose eval card runs live on /ask. Real client data only with written consent.
Open evals
Pass rate · gate · cost · incidents
Sample board · as of September 2026
Evals, or it didn’t happen.
A public board for AI in production: how often it passes its test set, the bar it has to clear, what a task costs, and when it last went wrong.
Systems on the board5
Above the gate3
Below the gate1
Real, not a sample1
sample Four fictional systems and one real one, Ask Voidgap.
01 · The board
Five systems, nothing hidden.
Hover or tab into a chart to read each run. The dashed line is the gate: below it, nothing ships.
Runs
Status
The real one · on this site
Ask Voidgap
in place
Answers questions from the pages of this website and cites them. Its eval card is public: eight questions, the page each answer must cite, one it must refuse. It runs on your account, in your browser, when you press the button.
Last incident
incident log · to decide
Last run
–/8yours, when you run it
Gate
to decide
Cost per task
to decide
No run history yet. Runs are not stored anywhere, so there is no line to draw. A public history needs storage and a decision: run history · to decide
Pass means Facts correct, right tone, no promise outside the policy
Last incident
Drafts quoted an outdated returns window after a policy change. Knowledge base updated; three cases added to the test set.
Data table
Run
Pass rate
vs gate
Note
21 Sep 2026
93.1 %
+3.1
14 Sep 2026
92.9 %
+2.9
7 Sep 2026
93.2 %
+3.2
31 Aug 2026
93.6 %
+3.6
24 Aug 2026
93.0 %
+3.0
17 Aug 2026
93.3 %
+3.3
10 Aug 2026
92.8 %
+2.8
3 Aug 2026
86.5 %
−3.5
Model version update scored below the gate. Not deployed.
27 Jul 2026
94.0 %
+4.0
20 Jul 2026
93.5 %
+3.5
13 Jul 2026
93.9 %
+3.9
6 Jul 2026
93.2 %
+3.2
Documents into the ERP
Delivery note extraction
sample
Last run
97.6 %above the gate
Gate
97 %520 test cases
Cost per task
€0.03per document
Pass means Every required field correct
Last incident
A scanner update rotated pages. Confidence dropped, 212 documents went to manual review, none reached the ERP unchecked.
Data table
Run
Pass rate
vs gate
Note
21 Sep 2026
97.6 %
+0.6
14 Sep 2026
97.5 %
+0.5
7 Sep 2026
97.1 %
+0.1
31 Aug 2026
97.4 %
+0.4
24 Aug 2026
97.2 %
+0.2
First run above the gate. Went live the week after.
17 Aug 2026
96.9 %
−0.1
10 Aug 2026
96.7 %
−0.3
3 Aug 2026
96.1 %
−0.9
27 Jul 2026
96.3 %
−0.7
20 Jul 2026
96.0 %
−1.0
13 Jul 2026
95.4 %
−1.6
6 Jul 2026
95.6 %
−1.4
Flags risky clauses for legal
Contract clause finder
sample
Last run
95.2 %below the gate
Gate
98 %150 test cases
Cost per task
€0.31per contract
Pass means Every clause on the checklist found, nothing invented
Last incident
None in production. It is not in production: below the gate it runs in shadow mode, next to the lawyers, not instead of them.
Data table
Run
Pass rate
vs gate
Note
21 Sep 2026
95.2 %
−2.8
14 Sep 2026
95.0 %
−3.0
7 Sep 2026
94.8 %
−3.2
31 Aug 2026
94.6 %
−3.4
24 Aug 2026
94.1 %
−3.9
17 Aug 2026
93.9 %
−4.1
10 Aug 2026
93.5 %
−4.5
3 Aug 2026
93.0 %
−5.0
27 Jul 2026
92.4 %
−5.6
20 Jul 2026
92.6 %
−5.4
13 Jul 2026
92.1 %
−5.9
6 Jul 2026
91.8 %
−6.2
Voice agent
Appointment booking by phone
sample
Last run
94.4 %above the gate
Gate
92 %260 test cases
Cost per task
€0.18per call
Pass means Right slot, right person, caller told it is an AI
Last incident
Callers with double-barrelled surnames were booked under half their name. 23 bookings corrected by hand; 40 name cases added to the test set.
Data table
Run
Pass rate
vs gate
Note
21 Sep 2026
94.4 %
+2.4
14 Sep 2026
94.8 %
+2.8
7 Sep 2026
94.6 %
+2.6
31 Aug 2026
94.3 %
+2.3
24 Aug 2026
94.7 %
+2.7
17 Aug 2026
94.9 %
+2.9
10 Aug 2026
94.2 %
+2.2
3 Aug 2026
94.6 %
+2.6
27 Jul 2026
94.8 %
+2.8
20 Jul 2026
94.3 %
+2.3
13 Jul 2026
94.5 %
+2.5
6 Jul 2026
94.0 %
+2.0
Weekly runs on a fixed test set. A run below the gate blocks the deploy; the line shows what was tested, not what went live. Hollow dots mark runs with a note.
02 · The four numbers
Four numbers per system.
If a vendor cannot give you these four, the system is a demo with a login.
1Pass rate
Share of the test set that passes on a run. The test set comes from real cases, with the expected result written down before the build.
2The gate
The pass rate the owner agreed to before launch. Below it, a change does not deploy. It is set per system, because 90 % is fine for a draft and not for a customs value.
3Cost per task
Model and infrastructure cost per finished task, not per token. The number finance can put next to the salary it saves or the hours it frees.
4Last incident
The last time it went wrong for a real user, with the date and what changed after. A board without incidents is either very new or not honest.