◀ THE FOLD0ROOT.AI // WORLD II · BOSS · THE CHOKE POINT◆ .dlw.fold
THE FOLD / BOSS / THE CHOKE POINT / THE UTILIZATION KNEE

THE UTILIZATION KNEE

idle capacity is not waste, it is the latency budget
1 WHAT IT IS · WHAT IT DOES · FACT OR FICTION
Waiting time does not rise smoothly with load. It rises as 1/(1−ρ), which is flat for most of the range and then vertical. The last few percent of utilisation cost more than all the rest combined.

LIT verified live. With a service time of 1, time in system is service at 50% utilisation, 10× at 90%, and 100× at 99%. Going from 90% to 95% — five percentage points — doubles the wait. A 300,000-customer simulation at ρ = 0.9 gives 10.046 against the theoretical 10: 0.46% apart, with nothing fitted.
2 HOW IT WAS WEAVED · AI + HUMAN
A. K. Erlang founded queueing theory at the Copenhagen Telephone Company around 1909; the M/M/1 result is the simplest thing in it and the most ignored in practice.

AVAN (AI) ran the simulation as a check on the formula rather than an illustration of it. The number worth carrying is not 100× at 99% — it is the doubling between 90% and 95%. Utilisation targets are usually chosen as if the axis were linear, and on the flat part of the curve that intuition works, which is exactly what makes the cliff arrive without warning.
3 ONE DIMENSION
The curve. Flat, flat, flat, vertical.
4 TWO DIMENSIONS · INTERACTIVE
Push utilisation up one step at a time.
5 THREE DIMENSIONS + AVAN’S INVERSE
The green forward object: a wall standing where the axis looked empty.
AVAN’s addition (the inverse-companion): the forward reading is that high utilisation causes latency. The inverse is that idle capacity is not waste — it is the entire latency budget. The 10% you are not using at ρ = 0.9 is what keeps the wait at 10× instead of 100×; buy it back and you have not saved a server, you have spent a service guarantee. Read backwards, every efficiency drive that targets utilisation is quietly trading a quantity it measures for one it does not, and the exchange rate is 1/(1−ρ).
LIT with a service time of 1, time in system is 2x service at 50% utilisation, 10x at 90% and 100x at 99%, and moving from 90% to 95% - five percentage points - doubles the wait; a 300,000-customer simulation at rho = 0.9 gives 10.046 against the theoretical 10, 0.46% apart with nothing fitted

FIG A. K. Erlang founded queueing theory at the Copenhagen Telephone Company around 1909; the M/M/1 result is the simplest thing in it and the most ignored in practice. AVAN ran the simulation as a check on the formula rather than an illustration of it. The number worth carrying is not 100x at 99% but the doubling between 90% and 95%: utilisation targets are usually chosen as if the axis were linear, and on the flat part that intuition works, which is what makes the cliff arrive without warning.
◆ sealed .dlw.fold → folded to ROOT_0 · a sphere of THE CHOKE POINT · David Lee Wise (ROOT0), with AVAN