pi0.5 — Spa-Bench Epoch 12 (Step 76,596)

This is the pi0.5 checkpoint evaluated in Spa-Bench, a real-robot benchmark of spatially grounded reasoning in vision-language-action policies.

Model details

Field Value
Model repository justintiensmith/pi05_Reasoning_Step_076596
Base model lerobot/pi05_base
Checkpoint End of epoch 12; step 76,596
Robot SO-101 single-arm manipulator
Inputs Fixed middle RGB, wrist RGB, six absolute joint positions, text instruction
Outputs Six absolute joint-position targets
Action horizon 50
Normalization Quantile normalization for state and action
Adaptation Vision-language backbone and action expert updated; no PEFT
Optimizer AdamW, peak learning rate 2.5e-5, weight decay 0.01
Schedule 1,000 warm-up updates, then cosine decay
Hardware and batch Four NVIDIA GH200 GPUs; 24 samples per device, global batch 96

The embedded train_config.json pins the training data to justintiensmith/VLA_Reasoning_Training_Dataset_1200@3c92bdd. That dataset contains 1,200 episodes, 612,733 frames, 321 instruction strings, and five RGB views; this policy consumes only the middle and wrist views. The historical run records the base-model identifier but not its immutable source revision. The current public base revision at archival review is b211f3d44c36b6acfcf7ae94a64e8e96f75a64ba and must not be assumed to be the unrecorded training revision.

The author-supplied original train_pi05_v3_isambard.sh launcher is archived with the thesis artifact. It records the Isambard environment, four-GPU launch, dataset pin, cache layout, checkpoint cadence, and resume behavior used for this run.

Physical evaluation

Condition Successes Rate
Familiar/in-distribution spatial instructions 82/120 68.3%
All withheld/OOD spatial configurations 122/300 40.7%
Matched OOD subset 54/120 45.0%
Matched direct-manipulation controls 100/120 83.3%

The matched-control gap was 38.3 percentage points. These are physical rollout results, not simulation metrics. The associated repository contains 920 recordings: episodes 0–899 are the scored protocol, while episodes 900–919 are additional Size Recognition recordings excluded from the reported analysis.

Intended use and limitations

This release supports reproduction and analysis of the Spa-Bench experiment. The results apply to this checkpoint, SO-101 embodiment, workspace, cameras, objects, and protocol. They do not establish general robot capability, and no held-out validation loss was used to select among checkpoints: epoch 12 was fixed before physical evaluation.

Robot policies can move hardware unexpectedly. Use conservative motion limits, an accessible emergency stop, a clear workspace, and direct supervision. Do not deploy this checkpoint for unattended or safety-critical operation.

The repository does not relicense third-party components; use is subject to the base model's Gemma terms and the licenses of the LeRobot software and released data.

Citation

Please cite the completed Spa-Bench MSc report, the thesis artifact, and the pi0.5 work referenced in the report.

Downloads last month
42
Safetensors
Model size
4B params
Tensor type
F32
·
BF16
·
Video Preview
loading

Model tree for justintiensmith/pi05_Reasoning_Step_076596

Finetuned
(680)
this model

Dataset used to train justintiensmith/pi05_Reasoning_Step_076596