# Evaluation protocol Exact-state replay on RoboCasa PickPlaceCounterToCabinet. Every run restores the same 160 saved scenes, so checkpoints are directly comparable. ## The episode set `pnpcountertocab_mimicgen8_exact160_replay_policyseed12345_20260506` — 160 episodes, **20 per object** for the 8 augmentation targets: ``` 0~ 19 donut/donut_5 80~ 99 steak/steak_8 20~ 39 jar/Jar023 100~119 syrup_bottle/SyrupBottle005 40~ 59 measuring_cup/MeasuringCup009 120~139 teapot/teapot_6 60~ 79 soap_dispenser/SoapDispenser010 140~159 teapot/teapot_7 ``` Drawn from a 1,160-episode base-60k rollout collection over 58 objects, keeping the episodes for these 8. That collection scored 45.0% overall (466/1035 completed); on these 160 it scored 15/160, so this is a deliberately hard slice — a policy doing well elsewhere still looks low here. Each episode ships `model_xml_gz`, `state_npz` and an `ep_meta` pickle; the wrapper restores the XML and the flattened MuJoCo state, so the initial condition is identical across runs rather than re-sampled. ## Settings — hold these fixed ```bash MODEL_PATH= # the only thing that should vary MYGROOT_ROOT= # the default path in the launcher is dead EVAL_CLIENT=<.../eval_robocasa_replay_state_grasp.py> REPLAY_STATE_ROOT=<.../pnpcountertocab_mimicgen8_exact160_replay_policyseed12345_20260506> N_EPISODES=160 EXPECTED_REPLAY_STATES=160 POLICY_SEED=12345 SEED_BASE=42 GPUS= ``` `POLICY_SEED` is the one that changes rollouts: it seeds the policy server's global RNG and the per-step action seed (`policy_seed + episode_index x stride + step`). It matters — on one checkpoint, 54321 gave 2/160 and 12345 gave 6/160, a 3x swing from the seed alone. **Do not read small differences between single-seed runs as a result.** `SEED_BASE` seeds the env (`SEED_BASE + worker_id`), but exact replay overwrites the scene immediately after reset, so it should not affect the initial state — unverified, so keep it at 42 anyway. Worker count follows the GPU list and episodes split evenly across workers, so the GPU count changes which worker seed each episode gets. Keep **four workers**. To put four workers on two cards, repeat them: `GPUS=2,3,2,3`. Each server needs roughly 20 GB, so two servers per card needs ~42 GB free — check before launching if training shares those cards, or the OOM takes the training down too. ## Which checkpoints Compare like with like. The runs completed so far: | model | checkpoint | steps | note | |---|---|---|---| | base | `pnpcountertocab_all502_gbs64_wandb_60k_save5k_.../checkpoint-60000` | 60k | the checkpoint every augmentation run fine-tunes from | | MimicGen | `mimicgen_natural_256_fixed/checkpoint-30000` | 30k | ~28 epochs over 256 episodes | | MimicGen (short) | `mimicgen_natural_256_short/checkpoint-1500` | 1.5k | same data and config, ~1.4 epochs | | VACE | `vace_batch64_30k_from_hf/checkpoint-30000` | 30k | | | actaug | `actaug_256/checkpoint-1500` | 1.5k | | Always evaluate the base checkpoint under the same harness rather than quoting a number from the run that produced the scenes. Those two are not the same measurement: the collection run recorded 15/160 while exact replay of the same episodes with the same policy seed gives **11/160**, so 6.9% is the baseline to compare against, not 9.4%. Step counts differ between runs, so a 1.5k checkpoint against a 30k one confounds training length with the data. Evaluate the same step count on both sides before attributing a difference to the dataset. ## Reading the output ``` // summary.json n_episodes, n_success, success_rate, and the full run config simulation_results.csv per episode: success, reward, policy_steps, video_path, layout_id, style_id, language, target_mesh_keys episodes.jsonl merged per-episode rows (the stage flags live here) workers/worker_NN_seed_S/ per-worker logs, videos, ep_meta pickles rollouts/episode_NNNNNN/ center.mp4 + wrist.mp4 + info.txt, after organize_videos.py ``` `policy_steps x action_horizon` equals `sim_step_budget` (47 x 16 = 752) on any episode that ran out the clock, which is what failure looks like — there is no early abort. ### Stage flags Stock RoboCasa success is `obj_inside_of(cab) and gripper_obj_far`: one boolean, which cannot say where a failure broke down. `eval_robocasa_replay_state_grasp.py` adds three per-episode flags, each latched if it ever held during the rollout: | flag | source | |---|---| | `grasped` | `env._check_grasp(gripper, object_geoms)` | | `lifted` | object body z rose more than 3 cm above its start height | | `in_cab` | `OU.obj_inside_of(env, "obj", env.cab)` | They localise the failure. On one run: 45 grasped, 29 lifted, 27 in cabinet, 25 succeeded — a third of successful grasps were dropped during the lift, while nearly everything lifted went on to finish. Without the flags all of that reads as a single success rate. ## Sampling a subset `--episode_indices` takes an arbitrary list, for a quick read before committing an hour per checkpoint. Five per object: ```bash python3 -c "print(','.join(str(b+i) for b in range(0,160,20) for i in range(5)))" # 0,1,2,3,4,20,21,22,23,24,40,...,144 ``` Per-object rates from 5 episodes are noisy; use it to decide whether a full run is worth it, not to compare checkpoints.