|
Download eval/EVALUATION.md from mlnha/pose6daug-scripts: direct link, hf CLI and curl.
- Browser
- Download file 5.5 kB
-
https://huggingface.co/mlnha/pose6daug-scripts/resolve/main/eval/EVALUATION.md
- Command line
-
hf download hf://mlnha/pose6daug-scripts/eval/EVALUATION.md
-
curl -L -o EVALUATION.md https://huggingface.co/mlnha/pose6daug-scripts/resolve/main/eval/EVALUATION.md
5.5 kB
| # Evaluation protocol | |
| Exact-state replay on RoboCasa PickPlaceCounterToCabinet. Every run restores the same 160 | |
| saved scenes, so checkpoints are directly comparable. | |
| ## The episode set | |
| `pnpcountertocab_mimicgen8_exact160_replay_policyseed12345_20260506` β 160 episodes, | |
| **20 per object** for the 8 augmentation targets: | |
| ``` | |
| 0~ 19 donut/donut_5 80~ 99 steak/steak_8 | |
| 20~ 39 jar/Jar023 100~119 syrup_bottle/SyrupBottle005 | |
| 40~ 59 measuring_cup/MeasuringCup009 120~139 teapot/teapot_6 | |
| 60~ 79 soap_dispenser/SoapDispenser010 140~159 teapot/teapot_7 | |
| ``` | |
| Drawn from a 1,160-episode base-60k rollout collection over 58 objects, keeping the episodes | |
| for these 8. That collection scored 45.0% overall (466/1035 completed); on these 160 it scored | |
| 15/160, so this is a deliberately hard slice β a policy doing well elsewhere still looks low | |
| here. Each episode ships `model_xml_gz`, `state_npz` and an `ep_meta` pickle; the wrapper | |
| restores the XML and the flattened MuJoCo state, so the initial condition is identical across | |
| runs rather than re-sampled. | |
| ## Settings β hold these fixed | |
| ```bash | |
| MODEL_PATH=<checkpoint> # the only thing that should vary | |
| MYGROOT_ROOT=<myGR00T tree> # the default path in the launcher is dead | |
| EVAL_CLIENT=<.../eval_robocasa_replay_state_grasp.py> | |
| REPLAY_STATE_ROOT=<.../pnpcountertocab_mimicgen8_exact160_replay_policyseed12345_20260506> | |
| N_EPISODES=160 EXPECTED_REPLAY_STATES=160 | |
| POLICY_SEED=12345 SEED_BASE=42 GPUS=<four entries> | |
| ``` | |
| `POLICY_SEED` is the one that changes rollouts: it seeds the policy server's global RNG and | |
| the per-step action seed (`policy_seed + episode_index x stride + step`). It matters β on one | |
| checkpoint, 54321 gave 2/160 and 12345 gave 6/160, a 3x swing from the seed alone. **Do not | |
| read small differences between single-seed runs as a result.** | |
| `SEED_BASE` seeds the env (`SEED_BASE + worker_id`), but exact replay overwrites the scene | |
| immediately after reset, so it should not affect the initial state β unverified, so keep it | |
| at 42 anyway. | |
| Worker count follows the GPU list and episodes split evenly across workers, so the GPU count | |
| changes which worker seed each episode gets. Keep **four workers**. To put four workers on two | |
| cards, repeat them: `GPUS=2,3,2,3`. Each server needs roughly 20 GB, so two servers per card | |
| needs ~42 GB free β check before launching if training shares those cards, or the OOM takes | |
| the training down too. | |
| ## Which checkpoints | |
| Compare like with like. The runs completed so far: | |
| | model | checkpoint | steps | note | | |
| |---|---|---|---| | |
| | base | `pnpcountertocab_all502_gbs64_wandb_60k_save5k_.../checkpoint-60000` | 60k | the checkpoint every augmentation run fine-tunes from | | |
| | MimicGen | `mimicgen_natural_256_fixed/checkpoint-30000` | 30k | ~28 epochs over 256 episodes | | |
| | MimicGen (short) | `mimicgen_natural_256_short/checkpoint-1500` | 1.5k | same data and config, ~1.4 epochs | | |
| | VACE | `vace_batch64_30k_from_hf/checkpoint-30000` | 30k | | | |
| | actaug | `actaug_256/checkpoint-1500` | 1.5k | | | |
| Always evaluate the base checkpoint under the same harness rather than quoting a number from | |
| the run that produced the scenes. Those two are not the same measurement: the collection run | |
| recorded 15/160 while exact replay of the same episodes with the same policy seed gives | |
| **11/160**, so 6.9% is the baseline to compare against, not 9.4%. | |
| Step counts differ between runs, so a 1.5k checkpoint against a 30k one confounds training | |
| length with the data. Evaluate the same step count on both sides before attributing a | |
| difference to the dataset. | |
| ## Reading the output | |
| ``` | |
| <EVAL_OUTPUT_BASE>/<EVAL_RUN_NAME>/ | |
| summary.json n_episodes, n_success, success_rate, and the full run config | |
| simulation_results.csv per episode: success, reward, policy_steps, video_path, | |
| layout_id, style_id, language, target_mesh_keys | |
| episodes.jsonl merged per-episode rows (the stage flags live here) | |
| workers/worker_NN_seed_S/ per-worker logs, videos, ep_meta pickles | |
| rollouts/episode_NNNNNN/ center.mp4 + wrist.mp4 + info.txt, after organize_videos.py | |
| ``` | |
| `policy_steps x action_horizon` equals `sim_step_budget` (47 x 16 = 752) on any episode that | |
| ran out the clock, which is what failure looks like β there is no early abort. | |
| ### Stage flags | |
| Stock RoboCasa success is `obj_inside_of(cab) and gripper_obj_far`: one boolean, which cannot | |
| say where a failure broke down. `eval_robocasa_replay_state_grasp.py` adds three per-episode | |
| flags, each latched if it ever held during the rollout: | |
| | flag | source | | |
| |---|---| | |
| | `grasped` | `env._check_grasp(gripper, object_geoms)` | | |
| | `lifted` | object body z rose more than 3 cm above its start height | | |
| | `in_cab` | `OU.obj_inside_of(env, "obj", env.cab)` | | |
| They localise the failure. On one run: 45 grasped, 29 lifted, 27 in cabinet, 25 succeeded β | |
| a third of successful grasps were dropped during the lift, while nearly everything lifted went | |
| on to finish. Without the flags all of that reads as a single success rate. | |
| ## Sampling a subset | |
| `--episode_indices` takes an arbitrary list, for a quick read before committing an hour per | |
| checkpoint. Five per object: | |
| ```bash | |
| python3 -c "print(','.join(str(b+i) for b in range(0,160,20) for i in range(5)))" | |
| # 0,1,2,3,4,20,21,22,23,24,40,...,144 | |
| ``` | |
| Per-object rates from 5 episodes are noisy; use it to decide whether a full run is worth it, | |
| not to compare checkpoints. | |