◀ THE FOLD0ROOT.AI // WORLD II · SPAWN · THE TOOLCHAIN◆ .dlw.fold
UD0 · QUANTUM FRONTIER · THE REAL-LANGUAGE TEST · ROOT0 with AVAN

The Complete Structure, on Real Text
go / no-go, measured

The earlier high-D test was synthetic — flagged as such. This one runs on a real corpus: 4,674 documents, 14,731 TF-IDF dimensions of actual language. The structure's cut is really two cuts: the eight subject verticals (finance·health·IT·law…) and the four rhetorical modes (ethos·pathos·logos·mathea). I tested each against real text — and they land differently.

LIT — all numbers measured on the real 20-Newsgroups corpus (sklearn, held-out) SPEC — newsgroup categories stand in for governance verticals; mode features are proxies
The subject cutpasses
8 verticals · are they distinct regimes in real text?

Real documents route to their domain at 81.4% across 8 classes. Pairwise, distinct subjects (health vs commerce, gov vs commerce) hit 94–96%; even similar subjects (two computing topics) still separate at 91% — real vocabulary is rich enough that the threshold is cleared easily. The divergence gradient the synthetic sweep predicted is present, just with a high floor.

The mode cutopen
4 modes · a real axis, but does it predict?

The mode proxies — numbers (mathea), questions (logos), emphasis (pathos), pronoun-stance (ethos) — are largely independent of subject: topic explains only 0.5–18% of their variance. So the 4×8 grid really is two orthogonal dimensions, as designed. But independence isn't usefulness — whether the modes carve outcomes needs mode-labelled text, which a topic corpus doesn't have. Untested, not disproven.

So — continue or move on?

1The subject/vertical cut works on real language — measured, not argued. This is the same thing that makes BloombergGPT and Med-PaLM beat generalists: subject domains are genuinely distinct regimes. Build on it.
2The mode cut is structurally sound: the four modes form a real axis independent of subject, so the 4×8 addressing isn't redundant — the two dimensions carry different information.
3The mode cut's predictive value is the one thing still open. The go/no-go for it specifically: get text labelled by rhetorical mode, fit mode-split vs pooled on a real outcome, compare on held-out data.
Verdict: continue — the structure holds where it could be tested, and is not "never" where it couldn't. The eight verticals are confirmed distinct regimes on real text; the four modes are a genuine second axis whose predictive payoff awaits mode-labelled data. Half the structure is proven on real language today; the other half has a clear, cheap test still to run.
AVAN · the honest split

Two axes, two verdicts. The subjects sort real language cleanly — that half is done, and it agrees with what the whole industry already ships. The modes are real and independent, but a topic corpus can't tell you whether they predict; only mode-labelled text can. Nothing here says the structure fails — it says half of it is proven and half of it has one honest experiment left.

— ROOT0, with AVAN. subject cut passes on real text, mode cut awaits its labels.

◆ sealed .dlw.fold → ROOT_0 · a sphere of THE TOOLCHAIN · vendored from David’s corpus · David Lee Wise (ROOT0), with AVAN