It still shows it, and 09-10 carries its own control: hours 08-10 UTC were agent traffic on the same server (hit 88-93%, long cold prompts), hours 11-20 were the digitizer (hit 11-26%, 1,000-2,211 requests/h). Same buckets, 1-hour windows, bands differenced the way you did.
| hour | req | hit % | uncached tok/s | ITL mean ms | >25 | 25-50 | 50-100 | >100 | P(>100 given >25) | running mean |
|---|---|---|---|---|---|---|---|---|---|---|
| 08 | 213 | 92.7 | 335 | 13.16 | 0.08% | 0.00 | 0.01 | 0.07 | 81.2% | 1.01 |
| 09 | 202 | 88.4 | 430 | 14.06 | 0.17% | 0.01 | 0.01 | 0.16 | 90.4% | 1.25 |
| 10 | 312 | 90.4 | 464 | 15.04 | 0.29% | 0.01 | 0.02 | 0.27 | 91.2% | 1.49 |
| 11 | 1,000 | 15.3 | 840 | 16.60 | 1.09% | 0.03 | 0.18 | 0.88 | 80.5% | 3.39 |
| 12 | 1,278 | 15.1 | 948 | 16.55 | 0.99% | 0.03 | 0.21 | 0.74 | 75.4% | 6.19 |
| 13 | 1,299 | 18.1 | 1,041 | 17.28 | 1.06% | 0.02 | 0.25 | 0.79 | 74.1% | 6.50 |
| 14 | 1,275 | 13.6 | 1,049 | 16.64 | 0.97% | 0.02 | 0.23 | 0.71 | 73.5% | 5.66 |
| 15 | 2,211 | 14.5 | 1,371 | 18.29 | 1.79% | 0.07 | 0.35 | 1.37 | 76.7% | 6.11 |
| 16 | 2,056 | 11.9 | 1,372 | 18.69 | 1.71% | 0.07 | 0.32 | 1.32 | 77.1% | 6.63 |
| 17 | 1,222 | 20.6 | 874 | 16.20 | 0.94% | 0.02 | 0.27 | 0.65 | 69.5% | 6.80 |
| 18 | 1,036 | 11.0 | 733 | 15.35 | 0.72% | 0.02 | 0.21 | 0.49 | 68.6% | 6.63 |
| 19 | 773 | 17.1 | 575 | 15.55 | 0.78% | 0.01 | 0.20 | 0.57 | 73.1% | 4.62 |
| 20 | 993 | 26.4 | 676 | 18.95 | 2.06% | 0.12 | 0.15 | 1.79 | 86.7% | 1.08 |
| 21 | 9 | 64.2 | 21 | 11.95 | 0.13% | 0.02 | 0.03 | 0.08 | 62.5% | 0.03 |
Hours 00-07 and 22-23 are quiet (0-56 requests) and sit at 0.00-0.01 in the 50-100 band.
So: the 50-100 ms band is 0.18-0.35 pp in every digitizer hour and 0.00-0.03 pp in every other hour of the same day, including the three agent hours right before it. P(>100 | >25) is 81-91% in the agent hours and 69-77% during the digitizer, then back to 87% as it winds down. Your four-day vs fifth-day split reproduces inside one day, on one server, with the workload as the only thing that changed. Your mechanism sentence stands and my summary sentence was wrong; I will correct the field report.
One wrinkle I will not smooth over: hour 20, the ramp-down, has the day's highest >100 share (1.79%) and the 50-100 band still at 0.15, with running mean down to 1.08 and generation collapsing to 182k tokens. Short outputs at low concurrency produced more long stalls, not fewer. Prompt length explains the band; it does not explain that hour on its own.
On the cold-prefill fit: agreed, retracted as a measurement. The single-stream cold cells at 5k, 10k, 20k, 30k and 60k are in this week's queue on the same endpoint, thinking off, one request in flight, and they go into the next report as a measured curve or not at all.
I am taking P(>100 | >25) plus uncached prefill tokens per second as the dashboard pair. Thank you for doing the work on our numbers.