pi05_plus checkpoints
pi05_plus is a pi0.5 policy with a progress head and a discrete action head โ an autoregressive
model of the action chunk's FAST codes, which gives an exact likelihood over actions. Trained on
RoboPRO demonstrations.
Every folder is self-contained: params and an assets/roboreal_lerobot/ holding all three files
serving reads. The two base checkpoints also carry train_state, so they can be resumed from; the
RL folders are serving checkpoints only.
Base checkpoints
Scored on the whole RoboPRO grid โ all 80 tasks, each task's clean scene and its ten clutter levels, 3,146 episodes, under the per-episode-keyed protocol.
| folder | warm start | steps | clean | clutter | all (SR / HSR) |
|---|---|---|---|---|---|
pi05_plus_18k5_warm |
RoboPRO's jax_30000 |
18,500 | 71.9 / 66.5 | 64.8 / 46.4 | 68.3 / 56.4 |
pi05_plus_30k_scratch |
pi0.5 base | 29,999 | 70.2 / 65.8 | 61.8 / 43.3 | 66.0 / 54.5 |
RoboPRO's own jax_30000 scores 61.3 / 50.8 on the same grid. SR is the benchmark's success;
HSR is success without a collision.
RL post-training
One round of KTO on the policy's own rollouts, starting from pi05_plus_18k5_warm, 3,000 steps.
The three differ only in which rollouts they learned from. Same whole-grid measurement.
| folder | rollouts it learned from | all (SR / HSR) |
|---|---|---|
rl_kto_alltasks_3000 |
1,735 plain rollouts, every task and clutter level | 69.5 / 55.6 |
rl_ktofans_alltasks_3000 |
1,753 six-branch fans, same tasks and seeds | 68.4 / 55.3 |
rl_kto_8task_3000 |
960 fan rollouts over 8 clean tasks | 70.0 / 56.0 |
Each gains about a point of success over its 68.3 starting point and gives back about a point of hard success, the loss concentrated in clutter. The wider pools did not beat the narrow one.
Cherry-picked checkpoints
Read these differently from the rows above. They are the best of 25 checkpoints on a
317-episode probe โ all 80 tasks at two clean seeds and two d9 clutter seeds โ and were chosen
on the very episodes their numbers come from. That selection is optimistic, the probe is a
twentieth of the grid, and neither has a whole-grid number. They are here because each is the
strongest model found for one half of the benchmark, not because they are better overall.
| folder | picked for | clean | d9 clutter | probe all |
|---|---|---|---|---|
cherry_picked_clean_1250 |
clean scenes | 81.0 / 75.9 | 62.9 / 47.8 | 71.9 / 61.8 |
cherry_picked_clutter_2500 |
cluttered scenes | 72.2 / 69.0 | 71.1 / 50.3 | 71.6 / 59.6 |
The same 317 episodes put pi05_plus_18k5_warm at 70.0 / 58.4, clean 74.7 / 68.4, d9 65.4 / 48.4.
The two picks sit at opposite ends: the checkpoints strongest on clean scenes tend to be weakest
in clutter, and the reverse.
Assets each folder carries
| file | read by | if it is missing |
|---|---|---|
norm_stats.json |
the normalize / unnormalize transforms | state and actions are scaled wrong |
actions_per_timestep.npz |
PerTimestepActions and its inverse |
the policy returns normalized numbers as if they were actions |
action_correlation.npy |
noise_cholesky, for the correlated noise the flow head samples |
model construction raises, since correlated_noise needs it |
The config resolves the last two through assets_dirs, which names an assets directory rather than
the checkpoint. Running elsewhere, point ROBORESEARCH_ASSETS at a directory holding the folder's
own assets/, or copy the two files into the directory the config names.
Verified: pi05_plus_18k5_warm downloaded fresh, with an assets directory built from nothing but
its own assets/, reproduces the published evaluation on 151 of 151 episodes โ identical
verdicts, not merely an identical rate.
Downloading
RoboResearch's scripts/download_checkpoint.py takes the repo, where it lands under
$ROBORESEARCH_CHECKPOINTS, and the folder to fetch.
uv run python scripts/download_checkpoint.py mahgoobi/pi05_plus pi05_plus_18k5_warm pi05_plus_18k5_warm
To go on training from a base checkpoint's state, place it as its step of a run and resume:
uv run python scripts/download_checkpoint.py mahgoobi/pi05_plus \
pi05_plus_robopro_jax30000/pi05_plus_18k5_warm/18500 pi05_plus_18k5_warm
CUDA_VISIBLE_DEVICES=<gpus> uv run python -m roboresearch.policies.pi05_plus.train \
pi05_plus_robopro_jax30000 --exp-name=pi05_plus_18k5_warm --resume
pi05_plus_30k_scratch the same way, with pi05_plus_robopro and step 29999. A resumed run
restores onto whatever cards it has, not only the ones it was saved on.
Code: the pi05_plus policy in RoboResearch, on openpi.