Rollout evaluation — task10_taskfinetune
Closed-loop rollout evaluation on task 10 = set_up_a_coffee_station_in_your_kitchen,
against the official BEHAVIOR-1K solution checkpoint for the same task.
This supersedes the "no rollout evaluation has been run" note in the model card.
Result
official checkpoint_1 |
task10_taskfinetune |
|
|---|---|---|
| Success rate | 5.0 % (1/20) | 0.0 % (0/20) |
| Avg q_score | 0.1750 | 0.1917 |
| Median q | 0.167 | 0.167 |
| Episodes at q = 0 | 8 | 7 |
| Best episode | 1.000 | 0.500 |
The fine-tune shows no reliable gain. It edges the official checkpoint on mean q (+0.0167) but loses the only solved episode in either arm. Paired sign test over the 20 instances: 8 better / 7 worse / 5 tied, p = 1.000 — indistinguishable from chance.
The single most informative row is instance 310: the official checkpoint scores 1.000 (the only solve anywhere in this study), the fine-tune 0.167. The fine-tune is marginally more consistent (7 zeros vs 8, never above 0.500); the official checkpoint is spikier and occasionally completes the task.
Per-instance q_score
| instance | official checkpoint_1 |
task10_taskfinetune |
delta |
|---|---|---|---|
| 301 | 0.167 | 0.500 | +0.333 |
| 302 | 0.167 | 0.000 | -0.167 |
| 303 | 0.000 | 0.333 | +0.333 |
| 304 | 0.000 | 0.000 | +0.000 |
| 305 | 0.000 | 0.000 | +0.000 |
| 306 | 0.167 | 0.000 | -0.167 |
| 307 | 0.500 | 0.333 | -0.167 |
| 308 | 0.000 | 0.167 | +0.167 |
| 309 | 0.000 | 0.167 | +0.167 |
| 310 | 1.000 | 0.167 | -0.833 |
| 311 | 0.000 | 0.500 | +0.500 |
| 312 | 0.000 | 0.333 | +0.333 |
| 313 | 0.333 | 0.000 | -0.333 |
| 314 | 0.167 | 0.167 | +0.000 |
| 315 | 0.167 | 0.500 | +0.333 |
| 316 | 0.333 | 0.167 | -0.167 |
| 317 | 0.000 | 0.000 | +0.000 |
| 318 | 0.167 | 0.000 | -0.167 |
| 319 | 0.167 | 0.333 | +0.167 |
| 320 | 0.167 | 0.167 | +0.000 |
Evaluation setup
Both arms ran on the same day, same nodes, same evaluator — the comparison is internally controlled.
| task | set_up_a_coffee_station_in_your_kitchen (task_id 10) |
| instances | 20, --mode public_test (ids 301–320) |
| horizon | 1.5x mean human-demo length = 9,400 steps (no --max-steps override) |
| control rate | 30 Hz -> 313 s = 5.2 min simulated per episode |
| runner | omnigibson.eval.eval_fast + chunk-aware policy server |
| render | trailing4 + flush3; per episode full=470, tail_render=1410, fast=8930 |
base_qvel frame |
robot (B1K_BASE_QVEL_FRAME=robot) |
| norm stats | qvelfix, sha256 ee2e1a97... — byte-identical between both arms |
| dtype | float32 |
| challenge year | 2026 |
| official baseline | IliaLarchenko/behavior_submission checkpoint_1, per b50_tasks.csv (10,set_up_a_coffee_station_in_your_kitchen,checkpoint_1), with the qvelfix norm-stats overlay so both arms share one input contract |
| evaluator build | eval_fast.py mtime 2026-08-14 16:47 |
Caveats
- n = 20 cannot resolve differences this small. Re-running the same checkpoint under an identical configuration has been observed to swing a single instance's q_score by 0.333, so run-to-run noise is comparable to the effect being measured. Repeats across seeds would be needed to separate 0.175 from 0.192.
- Not comparable to numbers produced before 2026-08-14.
eval_fast.pygained the trailing-render + annotator-flush behaviour on that date; its own source comment attributes a 3-5x q_score effect to the flush. Only runs sharing an evaluator build may be compared. - Step 29,999 is not necessarily the best checkpoint. Training used a flat 5e-6 LR with no decay phase, so there is no annealing into a minimum. Intermediate steps were retained.
q_scoreandsuccesscan disagree — an episode scoring q = 1.000 has been observed withsuccess = Falseon a related task. Report both.