fastwalker1118's picture
Add task-5 evaluation videos and per-instance results
6070718 verified
|
Raw
History Blame Contribute Delete
2.25 kB
# task-5 `setting_mousetraps` — evaluation videos
Rollout videos and per-instance scores from the stock BEHAVIOR v3.9.1 evaluator
(`omnigibson.eval.eval`, commit `9a03ca3f`), non-fast path, on the 20 official
`public_test` instances (ids 301–320).
## Layout
```
merged39999/ 20 videos + results.json mean q 0.6417 success 1/20
merged20000/ 20 videos + results.json mean q 0.6167 success 0/20
merged15000/ results.json only mean q 0.6250 success 2/20
```
Filenames encode the outcome: `inst<id>_idx<NN>_q<score>.mp4`.
Each video is h264 672x448 — head plus both wrist cameras tiled into one frame.
`merged15000` has no videos because that sweep predated `--write-video`.
## Eval configuration
Two settings differ from a stock submission run and both matter when reading these numbers:
- **`B1K_STATE_2026=1`** — the 61-dim `PROP_2026` state extraction. Without it the
2025 extractor's slices (`base_qvel` 253:256) run off the end of v3.9.x's 61-dim
proprioception and silently yield a 2-element state, so the policy receives a
constant input and scores 0. This flag is required, not optional.
- **correction rules disabled** (`--no-apply-eval-tricks`) — they index `state[14]`
and raise `IndexError` on the collapsed state; disabling them was the only way to
complete rollouts while diagnosing. Re-enabling is untested.
Also applied: 100-task / 1120-stage embedding tables, which these checkpoints
require and the 50-task solution repos do not provide.
## What the scores say
```
merged15000 0.6250 merged20000 0.6167 merged39999 0.6417
```
A 0.025 spread across 15k–40k training steps, identical medians (0.6667), and a
paired sign test of 39999 vs 15000 giving p = 0.754. There is no evidence these
checkpoints differ; training further did not help.
The consistent feature is a ceiling: **14 of 20 instances score exactly 0.6667**
and 19 of 20 run to the 15,295-step cap. `q_score` counts goal predicates that
flipped false→true divided by the total in that option, so a predicate already
true at reset counts in the denominator but can never earn credit. Whether the
0.6667 wall is an unachieved condition or an uncreditable pre-satisfied one is
not resolved by these files.