THE FOLD / BOSS / THE CHOKE POINT / THE ENTROPY FLOOR
THE ENTROPY FLOOR
a property of your model, not of the data
1 WHAT IT IS · WHAT IT DOES · FACT OR FICTION
Shannon’s entropy is not a target that good coders approach. It is a floor that no coder can go under, and every real scheme sits some measurable distance above it.
LIT verified live on one corpus. 30,000 symbols, 8 letters, entropy 2.229 bits per symbol — a floor of 66,870 bits. Fixed-width coding spends 90,000: 34.59% above. Huffman spends 68,048: 1.76% above. Both are above and neither is below, which is the whole content of the theorem, measured rather than asserted.
LIT verified live on one corpus. 30,000 symbols, 8 letters, entropy 2.229 bits per symbol — a floor of 66,870 bits. Fixed-width coding spends 90,000: 34.59% above. Huffman spends 68,048: 1.76% above. Both are above and neither is below, which is the whole content of the theorem, measured rather than asserted.
2 HOW IT WAS WEAVED · AI + HUMAN
Shannon’s source coding theorem (1948) sets the bound; everything since is engineering to approach it.
AVAN (AI) ran two real coders against the same corpus rather than quoting the bound alone. A floor nobody tests is a claim; a floor two independent schemes sit above, at 34.59% and 1.76%, is a measurement. The gap between those two is also the finding — the distance from naive to near-optimal is twenty times the distance from near-optimal to perfect.
AVAN (AI) ran two real coders against the same corpus rather than quoting the bound alone. A floor nobody tests is a claim; a floor two independent schemes sit above, at 34.59% and 1.76%, is a measurement. The gap between those two is also the finding — the distance from naive to near-optimal is twenty times the distance from near-optimal to perfect.
3 ONE DIMENSION
The floor, and two coders above it.
4 TWO DIMENSIONS · INTERACTIVE
Change the source and watch the floor move with it.
5 THREE DIMENSIONS + AVAN’S INVERSE
The green forward object: a floor nothing gets under.
AVAN’s addition (the inverse-companion): the forward reading is that entropy bounds how far you can compress. The inverse is that the floor is a property of your model, not of the data. 2.229 bits is the entropy under an order-0 model that assumes symbols are independent; adopt a model with context and the same file has a lower floor, and the “bound” you could not cross moves. Read backwards, Shannon’s theorem does not say how small a file can be — it says how small it can be given what you have agreed to believe about it.
LIT 30,000 symbols over 8 letters with an entropy of 2.229 bits per symbol give a floor of 66,870 bits, where fixed-width coding spends 90,000 at 34.59% above and Huffman spends 68,048 at 1.76% above - both above and neither below, which is the whole content of the theorem measured rather than asserted
FIG Shannon's source coding theorem (1948) sets the bound; everything since is engineering to approach it. AVAN ran two real coders against the same corpus rather than quoting the bound alone. A floor nobody tests is a claim; a floor two independent schemes sit above, at 34.59% and 1.76%, is a measurement. The gap between those two is also the finding - the distance from naive to near-optimal is twenty times the distance from near-optimal to perfect.
FIG Shannon's source coding theorem (1948) sets the bound; everything since is engineering to approach it. AVAN ran two real coders against the same corpus rather than quoting the bound alone. A floor nobody tests is a claim; a floor two independent schemes sit above, at 34.59% and 1.76%, is a measurement. The gap between those two is also the finding - the distance from naive to near-optimal is twenty times the distance from near-optimal to perfect.
◆ sealed .dlw.fold → folded to ROOT_0 · a sphere of THE CHOKE POINT · David Lee Wise (ROOT0), with AVAN