| case carried by | lines/char | full 54-way | identity (27) | case (2) |
|---|---|---|---|---|
| colour | 2.0 | 76.8% | 78.8% | 96.9% |
| an extra short line | 2.5 | 63.4% | 73.8% | 80.0% |
| capture | full 54-way | identity (27) | case (2) |
|---|---|---|---|
| colour preserved | 75.1% | 76.6% | 97.8% |
| greyscale collapse | 41.4% | 74.7% | 54.3% |
| channel jitter | case accuracy | identity |
|---|---|---|
| ±0% | 98.1% | 79.2% |
| ±10% | 97.4% | 76.9% |
| ±25% | 90.0% | 77.7% |
| ±40% | 81.1% | 75.2% |
Why colour is the right channel for case specifically. Case is one bit. Identity is 27 ways. Colour is affine-invariant — rotation, scale, shear and translation cannot touch it — so it is cheap but low-capacity, which makes it a bad channel for identity and a very good one for a binary. Geometry is the opposite: high-capacity, distortion-sensitive. Putting the one-bit distinction on the invariant channel and the 27-way distinction on the geometric one is the whole design, and the measurement says it works.
Your frequency point decides which case gets marked. Lowercase dominates running prose by roughly an order of magnitude, so lowercase is the unmarked form — plain ink, no colour decision, no second pigment. Uppercase pays. Same variable-length-code logic that put the single-line glyphs on the commonest letters. AMBER on the exact ratio: I did not measure case frequency, the corpus here is all lowercase.
One thing this does not do. It buys case for free but it does not raise the ceiling on identity, which sits at roughly 78% at severity 1.5 with a 140-unit MLP. That number is the recogniser, not the alphabet. Every line-count and channel experiment has now hit the same wall, which is a fairly strong hint that the next worthwhile move is a convolutional recogniser rather than another glyph variation.