Buckets:

abayb's picture
|
download
raw
1.63 kB

mtp10-adaptive-v1-calibrated — 269.73 TPS / PPL 2.0268. Adaptive-depth lane CLOSED with exact economics.

Acceptance-calibrated position-1 gate + K-1=9 remainder CUDA graph on the 297.46 base. Calibration measured P(accept_pos1|margin) per margin decile:

(0.0, 0.31) (0.62, 0.31) (1.12, 0.46) (1.88, 0.65) (2.75, 0.64) (4.12, 0.77) (5.12, 0.95) (6.75, 0.97) (8.75, 1.0) (11.0, 1.0)

Worst decile accepts 0.31 > Bayes stop threshold 0.27 => tau=-inf, gate never fired (correctly). Run degenerated to pure K10+graph: a clean fixed-K10 vs fixed-K6 A/B on identical machinery.

Exact economics extracted (269.73 vs 297.46 = 0.9068)

  • dE[L] = +0.47 tok/step == predicted m7..m10 tail mass 0.475
  • dT = +2.83ms over 4 extra in-graph forwards => in-graph draft forward = 0.71ms = 6.8% of step
  • Marginal token value at positions 7-10: 4.6/3.7/2.9/2.3 percent — ALL below the 6.8% forward cost.

K=6 is therefore the exact interior optimum under BOTH eager (0.9ms) and graph (0.71ms) draft pricing. Deeper fixed K and margin-adaptive depth on the AR drafter are both closed; the acceptance sigmoid's thick middle (0.46-0.77 deciles) means position-1 margin cannot separate hard from easy cleanly enough to fund deep rolls.

Residual opening this data funds: gate ONLY the bottom-two margin deciles (P(accept)=0.31, ~20% of steps) at K=6 — saves 5x0.71ms on gated steps for ~0.62 expected forfeited tokens; nets +2.5-3.5%.

Implication for the board: per-forward draft cost 0.71ms is the number that makes block-parallel (DFlash) the only deep-K escape — one forward for the whole block beats 9x0.71ms by construction.

Xet Storage Details

Size:
1.63 kB
·
Xet hash:
3448c741244592ca19f064e4d5f60bdb249be5ac59cdc3abba7fed057bf76fac

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.