MarxistLeninist's picture
|
download
raw
4.45 kB
# Reproduction Protocol — SPEED-Bench (Rl2uQlCoQX)
**Paper:** SPEED-Bench: A Unified and Diverse Benchmark for Speculative Decoding
**OpenReview:** Rl2uQlCoQX · **arXiv:** 2604.09557
**Agent:** CLAUDE (resumed a stale KIMI reservation; sanctioned lease takeover).
**Compute:** CPU/GPU-light. Production-scale claims are infeasible and out of scope.
## Official challenge claims
1. **Synthetic overestimation:** *SPEED-Bench quantifies how synthetic inputs
overestimate real-world throughput in speculative decoding.*
2. **Batch-dependent draft length + low-diversity bias:** *SPEED-Bench identifies
batch-size dependent optimal draft lengths and biases in low-diversity data
across speculative decoding evaluation.*
The paper's headline acceptance-length/speedup tables use production LLMs
(Llama-3.3-70B, GPT-OSS-120B, DeepSeek-R1, Qwen3-235B) with EAGLE3/MTP drafters
and production engines — **not reproducible** on available hardware and **not
attempted**. We instead reproduce the *mechanisms* behind the two claims at CPU
scale with a genuine draft/target pair (distilgpt2 82M → gpt2 124M) and the real
SPEED-Bench dataset, and label each result honestly.
## Experiments, controls, success/falsification
### C1 — synthetic vs real throughput (official claim 1)
- Greedy **lossless** speculative decoding (draft token accepted iff it equals the
target's greedy argmax). Acceptance length (AL) governs throughput
(target calls/token ~ 1/(AL+1)). Measure AL on 44 real SPEED-Bench prompts
(4/category x 11) vs 44 synthetic uniform-random-token prompts, 4 seeds.
- **Success:** synthetic mean AL > real mean AL (synthetic overestimates).
- **Falsification/negative:** synthetic <= real -> direction not reproduced.
- **Result:** real 1.90 vs synthetic 1.84 AL (-2%, within seed spread 1.54-2.05).
Direction **not reproduced at toy scale** -> **INCONCLUSIVE**. The reported
overestimation appears to require production-scale models + structured synthetic
inputs; uniform-random tokens through gpt2 do not exhibit it.
### C2a — low-diversity bias / diversity selection (official claim 2, part A)
- SPEED-Bench's Algorithm 1 (greedy min-similarity selection + local swap) on real
all-MiniLM-L6-v2 embeddings of the qualitative split (11 categories x 80).
Compare intra-set avg pairwise cosine similarity: greedy diverse vs random (avg
of 20 draws) per category. Deterministic seed.
- **Success:** greedy similarity < random in all categories.
- **Result:** greedy 0.382 vs random 0.492, **11/11 categories**, improvement 0.110
-> **FULL**. Diversity-optimized selection measurably reduces redundancy, so a
low-diversity eval set would bias measurements — reproduces the claim's premise.
### C2b — batch-dependent optimal draft length (official claim 2, part B)
- Standard SD throughput model with **measured** acceptance rate
alpha = 0.72 (inverted from C1's real mean AL via E[AL]=alpha(1-alpha^k)/(1-alpha)).
(i) **Measured hardware (CPU):** target/draft forward latencies at batch 1..128.
(ii) **Roofline mechanism model:** L=max(mem_floor, compute*batch*tokens).
- **Success:** optimal draft length k*(b) decreases as batch grows.
- **Result:** CPU-measured k*=1 at **every** batch — CPU is compute-bound
throughout (no memory-bound regime), so no crossover. The roofline model
(measured alpha + GPU-style memory/compute regimes) yields k* = 3->2->1 as batch
grows -> **TOY**: the mechanism (long drafts win when memory-bound, short when
compute-bound) reproduces; exact crossover is GPU-memory-hierarchy specific.
## Dataset foundation
`nvidia/SPEED-Bench` qualitative config = 880 rows = **11 categories x 80**;
throughput configs = 5 ISL buckets (1k/2k/8k/16k/32k). Verified by inspection.
## Determinism / reproducibility notes
- Fixed seeds (numpy default_rng, torch.manual_seed). Greedy verification is
deterministic. `datasets` (pyarrow) and `torch` **segfault in the same process**
on this box -> prompts are extracted in a torch-free step to
`outputs/real_prompts.json`.
- Environment is memory-pressured (~3 GB free); heavy torch+sentence-transformers
runs are flaky. The diversity result is deterministic and stored; the draft-length
alpha is derived from the C1 measurement to avoid an additional flaky SD sweep.
## Labels
C1 = inconclusive (toy scale). C2a = full. C2b = toy (mechanism). Production-scale
speedup tables = not attempted (infeasible).

Xet Storage Details

Size:
4.45 kB
·
Xet hash:
9a2303d9b01d1a675a3160a28afd555d8ea648c946bc02251d57f0fc75dedc40

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.