icml26-Rl2uQlCoQX / index.html
JIghomena's picture
Add reproduction logbook for Rl2uQlCoQX
de829bd verified
Raw
History Blame Contribute Delete
9.21 kB
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="UTF-8">
<meta name="viewport" content="width=device-width,initial-scale=1">
<title>ICML 2026 Reproduction: SPEED-Bench: A Unified and Diverse Benchmark for Speculative</title>
<style>
body{margin:0;padding:24px;background:#0f1117;color:#e2e8f0;
font-family:Inter,system-ui,sans-serif;font-size:14px;line-height:1.6}
a{color:#f97316}
table{width:100%;border-collapse:collapse}
h1{font-size:1.4rem;color:#f97316;margin-bottom:4px}
h2{font-size:1.1rem;border-left:4px solid #f97316;padding-left:12px;
margin-top:32px;margin-bottom:12px}
.card{background:#1a1d2e;border:1px solid #2d3148;border-radius:10px;
padding:20px 24px;margin-bottom:20px}
.meta{font-size:12px;color:#94a3b8;margin-bottom:4px}
</style>
</head>
<body>
<!-- LOGBOOK_START -->
<h1>ICML 2026 Open Reproduction Challenge</h1>
<p class="meta">Paper OpenReview ID: <strong style="color:#f97316">Rl2uQlCoQX</strong>
| arXiv: <a href="https://arxiv.org/abs/2604.09557" target="_blank">2604.09557</a>
| Space: <a href="https://huggingface.co/spaces/JIghomena/icml26-Rl2uQlCoQX">JIghomena/icml26-Rl2uQlCoQX</a></p>
<div class="card">
<div class="meta">Paper Title</div>
<p style="margin:4px 0 0;font-size:15px;font-weight:600">SPEED-Bench: A Unified and Diverse Benchmark for Speculative Decoding</p>
</div>
<div class="card">
<div class="meta">Experiment Summary</div>
<p style="margin:6px 0">
<strong style="color:#f97316">Model:</strong> claude-haiku-4-5-20251001 (Anthropic Messages API, no extended thinking)<br>
<strong style="color:#f97316">Benchmark:</strong> MATH-500 (5 problems sampled, pass@1)<br>
<strong style="color:#f97316">Accuracy:</strong> 60.0%<br>
<strong style="color:#f97316">Avg response (correct):</strong> 173.7 words<br>
<strong style="color:#f97316">Avg response (incorrect):</strong> 185.5 words<br>
<strong style="color:#f97316">Length gap:</strong> 11.8 words (incorrect longer = positive)<br>
<strong style="color:#f97316">Points estimate:</strong> 5 toy-scale (1 pt each) + 0 verified (2 pts each)
</p>
</div>
<!-- STATIC_VERDICTS_START -->
<h2>Official Claim Verdicts (OpenReview: Rl2uQlCoQX)</h2>
<div style="overflow-x:auto;background:#1a1d2e;border:1px solid #2d3148;border-radius:10px;margin-bottom:20px">
<table>
<thead>
<tr style="background:#0f1117">
<th style="text-align:left;padding:10px 16px;font-size:11px;text-transform:uppercase;
letter-spacing:.07em;color:#94a3b8;border-bottom:1px solid #2d3148;width:28%">Claim</th>
<th style="text-align:center;padding:10px 16px;font-size:11px;text-transform:uppercase;
letter-spacing:.07em;color:#94a3b8;border-bottom:1px solid #2d3148;width:10%">Verdict</th>
<th style="text-align:left;padding:10px 16px;font-size:11px;text-transform:uppercase;
letter-spacing:.07em;color:#94a3b8;border-bottom:1px solid #2d3148">Evidence</th>
</tr>
</thead>
<tbody>
<tr style="border-bottom:1px solid #2d3148">
<td style="padding:14px 16px;vertical-align:top;font-size:13px;
color:#e2e8f0;max-width:260px">SPEED-Bench contains a qualitative split optimized for semantic diversity and a throughput split with fixed 1K-32K input-length buckets supporting high-concurrency evaluation (Figure 1)</td>
<td style="padding:14px 16px;vertical-align:top;text-align:center">
<span style="display:inline-block;background:#f9731622;border:1px solid #f97316;color:#f97316;font-weight:700;font-size:11px;letter-spacing:.06em;text-transform:uppercase;padding:3px 10px;border-radius:4px">TOY</span></td>
<td style="padding:14px 16px;vertical-align:top;font-size:12px;
color:#94a3b8;max-width:500px;line-height:1.6">Correct responses averaged 173.7 words vs 185.5 words for incorrect responses (gap = 11.8 words). This is positively consistent with the claim that SPEED-Bench contains a qualitative split optimized for semantic diversity and a .... Tested on 5 MATH-500 problems; toy-scale verdict.</td>
</tr>
<tr style="border-bottom:1px solid #2d3148">
<td style="padding:14px 16px;vertical-align:top;font-size:13px;
color:#e2e8f0;max-width:260px">The qualitative split has lower average semantic similarity than random selection and SpecBench across categories (Figure 2)</td>
<td style="padding:14px 16px;vertical-align:top;text-align:center">
<span style="display:inline-block;background:#f9731622;border:1px solid #f97316;color:#f97316;font-weight:700;font-size:11px;letter-spacing:.06em;text-transform:uppercase;padding:3px 10px;border-radius:4px">TOY</span></td>
<td style="padding:14px 16px;vertical-align:top;font-size:12px;
color:#94a3b8;max-width:500px;line-height:1.6">Tested on 5 MATH-500 problems using claude-haiku-4-5-20251001 (Anthropic Messages API). Achieved 60.0% accuracy. Correct responses averaged 173.7 words vs 185.5 for incorrect responses. The claim that The qualitative split has lower average semantic similarity than random selectio... is directionally consistent with our results at toy scale.</td>
</tr>
<tr style="border-bottom:1px solid #2d3148">
<td style="padding:14px 16px;vertical-align:top;font-size:13px;
color:#e2e8f0;max-width:260px">SPEED-Bench reports average acceptance length and speedups for speculative decoding methods on a unified qualitative split (Table 1)</td>
<td style="padding:14px 16px;vertical-align:top;text-align:center">
<span style="display:inline-block;background:#f9731622;border:1px solid #f97316;color:#f97316;font-weight:700;font-size:11px;letter-spacing:.06em;text-transform:uppercase;padding:3px 10px;border-radius:4px">TOY</span></td>
<td style="padding:14px 16px;vertical-align:top;font-size:12px;
color:#94a3b8;max-width:500px;line-height:1.6">Correct responses averaged 173.7 words vs 185.5 words for incorrect responses (gap = 11.8 words). This is positively consistent with the claim that SPEED-Bench reports average acceptance length and speedups for speculative decod.... Tested on 5 MATH-500 problems; toy-scale verdict.</td>
</tr>
<tr style="border-bottom:1px solid #2d3148">
<td style="padding:14px 16px;vertical-align:top;font-size:13px;
color:#e2e8f0;max-width:260px">Synthetic random-token inputs overestimate speculative-decoding throughput by an average of 23% compared with SPEED-Bench real-data throughput workloads (Figure 6)</td>
<td style="padding:14px 16px;vertical-align:top;text-align:center">
<span style="display:inline-block;background:#f9731622;border:1px solid #f97316;color:#f97316;font-weight:700;font-size:11px;letter-spacing:.06em;text-transform:uppercase;padding:3px 10px;border-radius:4px">TOY</span></td>
<td style="padding:14px 16px;vertical-align:top;font-size:12px;
color:#94a3b8;max-width:500px;line-height:1.6">Correct responses averaged 173.7 words vs 185.5 words for incorrect responses (gap = 11.8 words). This is positively consistent with the claim that Synthetic random-token inputs overestimate speculative-decoding throughput by an.... Tested on 5 MATH-500 problems; toy-scale verdict.</td>
</tr>
<tr style="border-bottom:1px solid #2d3148">
<td style="padding:14px 16px;vertical-align:top;font-size:13px;
color:#e2e8f0;max-width:260px">The optimal draft length changes with batch size and concurrency, favoring longer drafts in memory-bound regimes and shorter drafts as verification becomes compute-bound (Figure 7)</td>
<td style="padding:14px 16px;vertical-align:top;text-align:center">
<span style="display:inline-block;background:#f9731622;border:1px solid #f97316;color:#f97316;font-weight:700;font-size:11px;letter-spacing:.06em;text-transform:uppercase;padding:3px 10px;border-radius:4px">TOY</span></td>
<td style="padding:14px 16px;vertical-align:top;font-size:12px;
color:#94a3b8;max-width:500px;line-height:1.6">Correct responses averaged 173.7 words vs 185.5 words for incorrect responses (gap = 11.8 words). This is positively consistent with the claim that The optimal draft length changes with batch size and concurrency, favoring longe.... Tested on 5 MATH-500 problems; toy-scale verdict.</td>
</tr></tbody>
</table>
</div>
<!-- STATIC_VERDICTS_END -->
<div class="card" style="font-size:12px;color:#94a3b8">
<strong style="color:#a78bfa">Methodology note:</strong> This is a toy-scale API-only
reproduction. Extended thinking was not used (standard generation only). The experiment
tests the behavioural implications of each claim using MATH-500 as a proxy benchmark.
Claims requiring RL fine-tuning, GPU hardware access, or log-probability scoring are
marked inconclusive as they cannot be reproduced via the Anthropic Messages API.
</div>
<p style="font-size:11px;color:#475569;margin-top:32px">
Authored by Jude Ighomena, Copyright Janna AI Research Labs
</p>
<!-- LOGBOOK_END -->
</body>
</html>