Add rollout eval results for task10_taskfinetune; retarget README to the new folder name
f3cb490 verified | # Rollout evaluation β `task10_taskfinetune` | |
| Closed-loop rollout evaluation on **task 10 = `set_up_a_coffee_station_in_your_kitchen`**, | |
| against the official BEHAVIOR-1K solution checkpoint for the same task. | |
| This supersedes the "no rollout evaluation has been run" note in the model card. | |
| ## Result | |
| | | official `checkpoint_1` | `task10_taskfinetune` | | |
| |---|---|---| | |
| | **Success rate** | **5.0 %** (1/20) | **0.0 %** (0/20) | | |
| | **Avg q_score** | 0.1750 | **0.1917** | | |
| | Median q | 0.167 | 0.167 | | |
| | Episodes at q = 0 | 8 | 7 | | |
| | Best episode | **1.000** | 0.500 | | |
| **The fine-tune shows no reliable gain.** It edges the official checkpoint on mean q | |
| (+0.0167) but loses the only solved episode in either arm. | |
| Paired sign test over the 20 instances: **8 better / 7 worse / 5 tied, p = 1.000** β | |
| indistinguishable from chance. | |
| The single most informative row is **instance 310**: the official checkpoint scores **1.000** | |
| (the only solve anywhere in this study), the fine-tune **0.167**. The fine-tune is marginally more | |
| consistent (7 zeros vs 8, never above 0.500); the official checkpoint is spikier and | |
| occasionally completes the task. | |
| ## Per-instance q_score | |
| | instance | official `checkpoint_1` | `task10_taskfinetune` | delta | | |
| |---|---|---|---| | |
| | 301 | 0.167 | 0.500 | +0.333 | | |
| | 302 | 0.167 | 0.000 | -0.167 | | |
| | 303 | 0.000 | 0.333 | +0.333 | | |
| | 304 | 0.000 | 0.000 | +0.000 | | |
| | 305 | 0.000 | 0.000 | +0.000 | | |
| | 306 | 0.167 | 0.000 | -0.167 | | |
| | 307 | 0.500 | 0.333 | -0.167 | | |
| | 308 | 0.000 | 0.167 | +0.167 | | |
| | 309 | 0.000 | 0.167 | +0.167 | | |
| | 310 | 1.000 | 0.167 | -0.833 | | |
| | 311 | 0.000 | 0.500 | +0.500 | | |
| | 312 | 0.000 | 0.333 | +0.333 | | |
| | 313 | 0.333 | 0.000 | -0.333 | | |
| | 314 | 0.167 | 0.167 | +0.000 | | |
| | 315 | 0.167 | 0.500 | +0.333 | | |
| | 316 | 0.333 | 0.167 | -0.167 | | |
| | 317 | 0.000 | 0.000 | +0.000 | | |
| | 318 | 0.167 | 0.000 | -0.167 | | |
| | 319 | 0.167 | 0.333 | +0.167 | | |
| | 320 | 0.167 | 0.167 | +0.000 | | |
| ## Evaluation setup | |
| Both arms ran on the same day, same nodes, same evaluator β the comparison is internally controlled. | |
| | | | | |
| |---|---| | |
| | task | `set_up_a_coffee_station_in_your_kitchen` (task_id 10) | | |
| | instances | 20, `--mode public_test` (ids 301β320) | | |
| | horizon | 1.5x mean human-demo length = **9,400 steps** (no `--max-steps` override) | | |
| | control rate | 30 Hz -> 313 s = 5.2 min simulated per episode | | |
| | runner | `omnigibson.eval.eval_fast` + chunk-aware policy server | | |
| | render | trailing4 + flush3; per episode `full=470`, `tail_render=1410`, `fast=8930` | | |
| | `base_qvel` frame | **robot** (`B1K_BASE_QVEL_FRAME=robot`) | | |
| | norm stats | qvelfix, sha256 `ee2e1a97...` β **byte-identical between both arms** | | |
| | dtype | float32 | | |
| | challenge year | 2026 | | |
| | official baseline | `IliaLarchenko/behavior_submission` **checkpoint_1**, per `b50_tasks.csv` (`10,set_up_a_coffee_station_in_your_kitchen,checkpoint_1`), with the qvelfix norm-stats overlay so both arms share one input contract | | |
| | evaluator build | `eval_fast.py` mtime 2026-08-14 16:47 | | |
| ## Caveats | |
| 1. **n = 20 cannot resolve differences this small.** Re-running the *same* checkpoint under an | |
| identical configuration has been observed to swing a single instance's q_score by 0.333, so | |
| run-to-run noise is comparable to the effect being measured. Repeats across seeds would be | |
| needed to separate 0.175 from 0.192. | |
| 2. **Not comparable to numbers produced before 2026-08-14.** `eval_fast.py` gained the | |
| trailing-render + annotator-flush behaviour on that date; its own source comment attributes a | |
| 3-5x q_score effect to the flush. Only runs sharing an evaluator build may be compared. | |
| 3. **Step 29,999 is not necessarily the best checkpoint.** Training used a flat 5e-6 LR with no | |
| decay phase, so there is no annealing into a minimum. Intermediate steps were retained. | |
| 4. `q_score` and `success` can disagree β an episode scoring q = 1.000 has been observed with | |
| `success = False` on a related task. Report both. | |