◀ THE FOLD0ROOT.AI // WORLD II · RESPAWN · THE RESURRECT◆ .dlw.fold
THE FOLD / RESPAWN / THE RESURRECT / THE ONLY-FAILED PROBE

THE ONLY-FAILED PROBE

a detector nobody ever made say yes
1 WHAT IT IS · WHAT IT DOES · FACT OR FICTION
A detector that has only ever caught things has not been tested. Every catch is evidence about its recall and none at all about its false-positive rate — and two detectors with identical recall can differ enormously on how often they fire at nothing. On a validation set made only of real defects they are indistinguishable. The fix is not more positives; it is negatives, and the arithmetic of how many you need is unforgiving.

LIT verified live: two detectors with identical recall catch 4,737 and 4,744 of 5,000 real defects — statistically the same instrument; yet one fires on 2.0% of clean cases and the other on 60.1%, a difference no amount of positive testing could reveal; k clean controls that all pass bound the false-positive rate at 3/k by the rule of three, giving <30.0% at k=10 and <0.3% at k=1000; and with zero negative controls the bound is infinite.
2 HOW IT WAS WEAVED · AI + HUMAN
David (human) wrote the problem into Track C as a question rather than a claim — does the probe deserve to be trusted? — and answered it by exercising the PASS path and all seven exit paths end to end. Seated at THE RESURRECT: the probe only becomes credible once it has been made to say yes.

AVAN (AI) put the rule of three on the page because it makes the cost visible. Ten clean controls sound like diligence and bound the false-positive rate only below 30%; getting under 1% takes three hundred. That is why positives-only validation is so common — not carelessness, but because the negatives are expensive and produce nothing exciting when they pass. The asymmetry is worth naming plainly: a catch is a story and a clean pass is a line in a log, and the second one is what actually bounds the instrument.
3 ONE DIMENSION
Two detectors, identical where you looked, unrecognisable where you did not.
4 TWO DIMENSIONS · INTERACTIVE
Add clean controls and watch the bound come down, slowly.
5 THREE DIMENSIONS + AVAN’S INVERSE
The green forward object: the half of the space that was measured, and the half that was not.
AVAN’s addition (the inverse-companion): the forward reading is “test on negatives too.” The inverse is that a detector’s reputation is built entirely out of its true positives, because those are the only outcomes anyone narrates. Nobody writes up the morning the gate stayed quiet. So the evidence that reaches a decision-maker is systematically the half that cannot bound the false-positive rate, and the instrument looks better the more it fires. Read backwards, an instrument with a memorable track record is one whose weakest property has never been measured, and the fix is to make the quiet passes countable.
LIT two detectors with identical recall catch 4,737 and 4,744 of 5,000 real defects, statistically the same instrument; yet one fires on 2.0% of clean cases and the other on 60.1%, a difference no amount of positive testing could reveal; k clean controls that all pass bound the false-positive rate at 3/k by the rule of three, giving <30.0% at k=10 and <0.3% at k=1000; and with zero negative controls the bound is infinite

FIG The rule of three makes the cost visible: ten clean controls sound like diligence and bound the false-positive rate only below 30%; getting under 1% takes three hundred. That is why positives-only validation is so common — not carelessness, but because negatives are expensive and produce nothing exciting when they pass. A catch is a story and a clean pass is a line in a log, and the second is what actually bounds the instrument.
◆ sealed .dlw.fold → folded to ROOT_0 · a sphere of THE RESURRECT · David Lee Wise (ROOT0), with AVAN