THE FOLD / RESPAWN / THE PHOENIX / THE PREDICATE
THE PREDICATE
what it automates is the re-running, not the judgement
1 WHAT IT IS · WHAT IT DOES · FACT OR FICTION
claimlink states its limit before its features: this tool does not read prose and decide what it means. A human writes each claim down beside an executable predicate, and claimlink runs the predicate and reports. What it automates is the re-running, not the judgement. That splits total error cleanly in two. Predicate error is the human’s — the predicate does not capture what the sentence actually claims. Staleness error is the machine’s — the claim broke and nobody has looked since. Running more often drives the second toward zero and leaves the first exactly where it was.
LIT verified live: with an 8% predicate error and a 0.4% per-commit invalidation rate, re-running at cadence k gives predicate 7.8% + stale 0.2% at k=1, rising to predicate 8.1% + stale 56.9% at k=1000; staleness rises as the cadence loosens while predicate error does not move at all; re-running on every commit drives staleness to 0.17% of elapsed time, leaving 7.9% total — the predicate error and nothing else; and total error never falls below the predicate error at any cadence.
LIT verified live: with an 8% predicate error and a 0.4% per-commit invalidation rate, re-running at cadence k gives predicate 7.8% + stale 0.2% at k=1, rising to predicate 8.1% + stale 56.9% at k=1000; staleness rises as the cadence loosens while predicate error does not move at all; re-running on every commit drives staleness to 0.17% of elapsed time, leaving 7.9% total — the predicate error and nothing else; and total error never falls below the predicate error at any cadence.
2 HOW IT WAS WEAVED · AI + HUMAN
David (human) put the limitation first in the docstring, which is a choice with consequences — it makes the tool harder to oversell, including by its own author. Seated at THE PHOENIX: re-running revives the claim’s freshness, over and over, and never once touches whether the predicate was right.
AVAN (AI) measured the decomposition rather than asserting it, and the useful number is the floor. At every cadence tested, total error is at least the predicate error; the two components do not trade against each other, so no amount of automation substitutes for the judgement that went into writing the predicate. That is the honest ceiling on what this class of tool can do, and it is worth knowing before adopting one: a claim checker cannot make claims true, and cannot even make them checked in the sense people usually mean. It can only guarantee that whatever check a human already wrote has been run recently. Which is a real and unglamorous thing to guarantee, and it is exactly what the docstring says it is.
AVAN (AI) measured the decomposition rather than asserting it, and the useful number is the floor. At every cadence tested, total error is at least the predicate error; the two components do not trade against each other, so no amount of automation substitutes for the judgement that went into writing the predicate. That is the honest ceiling on what this class of tool can do, and it is worth knowing before adopting one: a claim checker cannot make claims true, and cannot even make them checked in the sense people usually mean. It can only guarantee that whatever check a human already wrote has been run recently. Which is a real and unglamorous thing to guarantee, and it is exactly what the docstring says it is.
3 ONE DIMENSION
Total error, split. One half responds to cadence. The other is flat.
4 TWO DIMENSIONS · INTERACTIVE
Turn the cadence up as far as you like and watch where it stops helping.
5 THREE DIMENSIONS + AVAN’S INVERSE
The green forward object: two error surfaces, only one of them tilted.
AVAN’s addition (the inverse-companion): the forward reading is “automation has a floor.” The inverse is that the floor is where the tool puts the human back, deliberately and in a named place. A system that hid the predicate — that read the prose and decided for itself — would not have a lower error, it would have an unlocatable one, distributed through a model nobody can point at. Writing the predicate by hand keeps the judgement in a file, with a line number, arguable. Read backwards, the honest limit is not a shortcoming of the design; it is the design, and the 7.9% floor is legible precisely because somebody chose to leave it visible.
LIT with an 8% predicate error and a 0.4% per-commit invalidation rate, re-running at cadence k gives predicate 7.8% + stale 0.2% at k=1, rising to predicate 8.1% + stale 56.9% at k=1000; staleness rises as the cadence loosens while predicate error does not move at all; re-running on every commit drives staleness to 0.17% of elapsed time, leaving 7.9% total โ the predicate error and nothing else; and total error never falls below the predicate error at any cadence
FIG The useful number is the FLOOR. At every cadence tested, total error is at least the predicate error; the two components do not trade against each other, so no amount of automation substitutes for the judgement that wrote the predicate. That is the honest ceiling on this class of tool: it cannot make claims true, and cannot even make them checked in the sense people usually mean โ only guarantee that whatever check a human already wrote has been run recently.
FIG The useful number is the FLOOR. At every cadence tested, total error is at least the predicate error; the two components do not trade against each other, so no amount of automation substitutes for the judgement that wrote the predicate. That is the honest ceiling on this class of tool: it cannot make claims true, and cannot even make them checked in the sense people usually mean โ only guarantee that whatever check a human already wrote has been run recently.
◆ sealed .dlw.fold → folded to ROOT_0 · a sphere of THE PHOENIX · David Lee Wise (ROOT0), with AVAN