The Lab · field evidence, not marketing
An AI judging another AI is not verification.
It is two systems agreeing on a story, and agreement is cheap. We spent this month trying to break that claim four different ways. It held three times. The fourth time, we were the ones who broke.
Every agentic AI stack now ships some version of the same failure: one model reading another model’s output and calling that a check. It isn’t. A judge reading a candidate is reading its confidence, its polish, its shape, not a fact about whether it is right. This is not a hypothetical. It is measured, published, and reproduced below, including once against us.
The Alibi Gate
A liar does not merely have to lie. He has to manufacture the evidence that covers it. Trust is not asserted here. It is priced, in the same currency as the thing being claimed.
Commit Before You Judge
A published paper drove an AI judge's false-positive rate from 72% to under 2% with one change: make it commit before it looks. We tried to reproduce the failure that fixes. Twice, honestly, nothing. Then we became the failure ourselves.
0.72 → 0.94 pass rate, 0.20 true accuracy
The Butterfly Budget
Two numbers, a trillionth apart. Same rule, sixty steps each. Drag the slider, hit run, watch a difference too small to matter become the entire answer.
10⁹¹³× amplification, 60 steps
Speak First
Same judge. Same candidate answer. One clock moved. Watch a 0.719 false-positive rate collapse to 0.012 in real time, live, in front of you.
0.719 → 0.012, one timing change
Two papers back every claim on this page. Nothing here is illustration standing in for proof; the proof came first, these are what happened when we went looking for where it could be broken. Read the math, or don’t, and lose the argument to whoever did.