codex-benchmark / index.html
yzzhao's picture
Publish codex SIMPLE seed-0 evaluation evidence
c465e3f verified
Raw History Blame Contribute Delete
4.45 kB
<!doctype html>
<html lang="en"><head><meta charset="utf-8"><meta name="viewport" content="width=device-width,initial-scale=1"><meta name="description" content="Codex robotics benchmark: original and revised instructions, seed 0, GPT-6 Astra high. Explore task videos, sessions and evidence across simulator families, including SIMPLE G1."><title>Codex Benchmark · RLE-Bench</title><link rel="icon" href="favicon.svg"><link rel="stylesheet" href="style.css"><script defer src="app.js"></script></head>
<body><header class="hero"><div class="wrap"><div class="topbar"><a class="brand" href="#"><span class="brand-icon" aria-hidden="true">◇</span> RLE-Bench</a><nav aria-label="Primary"><a href="#tasks">Tasks</a><a href="https://huggingface.co/datasets/RLE-Bench/codex-benchmark" target="_blank" rel="noopener">Dataset ↗</a><a href="REPORT.md" target="_blank" rel="noopener">Report ↗</a></nav></div><div class="hero-content"><div class="eyebrow">ROBOTICS EVALUATION / STOCK CODEX</div><h1>Codex Benchmark</h1><p>GPT-6 Astra · high · seed 0 · original &amp; modified instructions</p><div class="hero-bottom"><span>84 selected tasks across 5 simulator families. Native success criteria.</span><a class="hero-button" href="#task04">Explore the results <span>↗</span></a></div></div></div></header>
<main class="wrap"><div id="loading" role="status">Loading benchmark…</div><div id="browse" hidden><div class="edition"><span id="edition-count">Loading results…</span><span class="chip" id="edition-label">Benchmark results</span></div><section aria-labelledby="overview-heading"><div class="section-top"><h2 id="overview-heading">Task overview</h2><a href="summary.json" target="_blank" rel="noopener">Summary JSON ↗</a></div><div class="table-wrap"><table><thead><tr><th>Environment</th><th>Coverage</th><th>Native success</th><th>Input tokens</th><th>Output tokens</th><th>Usage coverage</th></tr></thead><tbody id="overview"></tbody></table></div><p class="footnote">Current recorded result per task · Seed 0 · Native success criteria. Each task shows its latest result and the instruction used for that recording. Rates include valid completed results only; pending tasks are excluded. No pooled score across environments. Input includes cached tokens; output includes reasoning tokens.</p></section><section id="failure-overview" lang="zh-CN" aria-label="失败原因"></section><p id="readiness-note" class="readiness-note" hidden></p><div id="tasks" class="task-controls"><nav id="family-nav" aria-label="Task families"></nav><div class="filters"><label class="search-label"><span class="sr-only">Search tasks</span><input id="search" type="search" placeholder="Search 84 tasks…"></label><label><span class="sr-only">Filter by status</span><select id="status-filter"><option value="all">All tasks</option><option value="completed">Results available</option><option value="pending">Pending</option></select></label><label><span class="sr-only">Filter by failure category</span><select id="failure-filter" aria-label="Filter by failure category"><option value="all">All failure categories</option></select></label><span id="filter-count" aria-live="polite"></span></div></div><div id="families"></div><p id="empty" hidden>No tasks match this search.</p><section class="protocol-box"><h2>The evaluation protocol</h2><div class="protocol-grid"><div><b>Instruction versions</b><p>Modified instructions are labeled, with changes highlighted. Each recording retains the instruction it used; native success predicates are unchanged.</p></div><div><b>Single seed</b><p>One independent episode, seed 0. GPT-6 Astra with high reasoning effort. Stock Codex CLI 0.160.0; each task starts with a fresh workspace and no inherited trajectory.</p></div><div><b>Native control budgets</b><p>RoboCasa and LIBERO: 6,000 steps at 20 Hz. RoboTwin and RoboDojo: 7,500 frames at 25 Hz. SIMPLE G1: 10,000 or 15,000 control frames at 50 Hz; see each task protocol. Stepped simulation, with an eight-hour wall-clock limit.</p></div></div></section></div><article id="detail" hidden></article></main>
<footer class="wrap"><span>RLE-Bench · Codex Benchmark</span><nav aria-label="Downloads"><a href="data.json">Data JSON</a><a href="episodes.csv">CSV</a><a href="DATA_FORMAT.md">Data format</a><a href="THIRD_PARTY_NOTICES.md">Attribution</a><a href="https://huggingface.co/datasets/RLE-Bench/codex-benchmark">Dataset ↗</a></nav></footer></body></html>