◀ THE FOLD0ROOT.AI // WORLD II · GLITCH · RACE CONDITION◆ .dlw.fold
THE FOLD / GLITCH / RACE CONDITION / THE ZIPF

THE ZIPF

a law that is not evidence
1 WHAT IT IS · WHAT IT DOES · FACT OR FICTION
Rank the words of any text by frequency and the counts fall off like 1/rank. Zipf’s law turns up in language, city sizes, income, web traffic — and it has been taken as evidence of deep organising principles for eighty years. In 1957 George Miller pointed out the problem: a monkey hitting random keys, including a space bar, produces Zipf’s law too. The law is not a fingerprint of meaning. It is what you get from any process that makes short things common and long things rare.

LIT verified live on 900,000 random keystrokes: 84,398 distinct “words” from 132,608 tokens, with a rank-frequency exponent of 0.899. The mechanism is exact — every word of length L has the same probability, so the mean count over all ML possible words falls by a factor of 0.031688, 0.031205, 0.032023, 0.031017 against a predicted (1−p)/M = 0.031538; and the fraction of possible words actually seen tracks the Poisson prediction 1−e−m at every length — 59.63% against 59.50%, 2.859% against 2.853%, 0.0897% against 0.0897%. A uniform-word control gives 0.056, no power law at all.
2 HOW IT WAS WEAVED · AI + HUMAN
David (human) seated this at RACE CONDITION: order that looks designed, arriving out of nothing but unsynchronised chance.

AVAN (AI) wrote a claim that the sweep destroyed. The first version asserted that every possible word of length ≤ 3 appears — and only 10,480 of 17,576 do. The right response was not to raise the sample size until the sentence became true, but to notice that the shortfall is exactly Poisson: with mean count m, the fraction seen should be 1−e−m, and it is, to two decimal places at every length. A second error followed immediately: the per-letter ratio was computed by averaging over observed words, which truncates at 1 and gave 0.329 at length 4 against a predicted 0.0315. Averaged over all ML possible words the ratio is right at every length. Both mistakes were the same mistake — conditioning on having seen something, and then measuring. One figure needs stating plainly: the branching structure gives an analytic exponent of −log((1−s)/M)/log M = 1.061, and the fitted value is 0.899, about 15% below it. That is not a defect in the theory or the fit — the regression window covers only the first few word lengths, and beyond length 3 the tail is so undersampled that it flattens. The Heaps exponent, measured over the same text, comes out 0.9445 against a predicted 1/1.061 = 0.9426, agreeing far better. Reporting the fitted 0.899 as though it confirmed the analytic 1.061 would have been the third version of the same error.
3 ONE DIMENSION
Rank against frequency, log-log. Random typing, and a control that has no law at all.
4 TWO DIMENSIONS · INTERACTIVE
The staircase underneath the law: one step per word length.
5 THREE DIMENSIONS + AVAN’S INVERSE
The green forward object: the word tree, each level M times wider and (1−p)/M times rarer.
AVAN’s addition (the inverse-companion): the forward reading is “Zipf’s law is not evidence of linguistic structure.” The inverse is that the law is a fact about the alphabet, not about the writer. Branching multiplies the count of words at each length by M and divides their probability by M/(1−p), and the rank-frequency curve is just those two exponentials plotted against each other. Nothing in the derivation knows what a word means. Read backwards, this is the general hazard of shape-matching evidence: a distribution that many mechanisms produce cannot discriminate between them, and the more universal a law looks, the less any single sighting of it tells you.
LIT on 900,000 random keystrokes: 84,398 distinct 'words' from 132,608 tokens, with a rank-frequency exponent of 0.899; every word of length L has the same probability, so the mean count over ALL M^L possible words falls by 0.031688, 0.031205, 0.032023, 0.031017 against a predicted (1-p)/M = 0.031538; the fraction of possible words actually seen tracks the Poisson prediction 1-e^-m at every length (59.63% vs 59.50, 2.859% vs 2.853, 0.0897% vs 0.0897); and a uniform-word control gives 0.056, no power law at all

FIG A claim was written and the sweep destroyed it. The first version asserted that EVERY possible word of length <= 3 appears - only 10,480 of 17,576 do. The right response was not to raise the sample size until the sentence became true, but to notice the shortfall is exactly Poisson: with mean count m the fraction seen should be 1-e^-m, and it is, to two decimals at every length. A second error followed at once: the per-letter ratio was computed over OBSERVED words, which truncates at 1 and gave 0.329 at length 4 against a predicted 0.0315. Both mistakes were the same mistake - conditioning on having seen something, then measuring. One figure needs stating plainly: the branching structure gives an ANALYTIC exponent of -log((1-s)/M)/log M = 1.061, and the fitted value is 0.899, about 15% below it. The regression window covers only the first few word lengths and beyond length 3 the tail is so undersampled it flattens. The Heaps exponent comes out 0.9445 against a predicted 1/1.061 = 0.9426, agreeing far better. Reporting 0.899 as though it confirmed 1.061 would have been the third version of the same error. Miller made the point in 1957.
◆ sealed .dlw.fold → folded to ROOT_0 · a sphere of RACE CONDITION · David Lee Wise (ROOT0), with AVAN