MorphAct Pi0.5: PiperX Put Drawer
An intermediate MorphAct adapter trained on the PiperX put-cube-in-drawer dataset. This release is optimizer step 14,000 of a planned 50,000-step run. It has not been evaluated on the real robot; no task success rate is claimed.
Contents and loading
adapter.safetensors contains the trained MorphAct generators and trainable vision/interface parameters. metadata.json preserves the adapter architecture and target manifest; the dataset normalization statistics and descriptive inference_config.json preserve preprocessing details.
Use the matching MorphAct Pi0.5 PyTorch adapter loader with the pi05_piperx_put_drawer config, or an equivalent config implementing the supplied inference contract. The training-specific Python config is not included in this weight repository. This checkpoint is an adapter rather than a standalone full base model.
Obtain an OpenPI-compatible Pi0.5 PyTorch base separately and supply its local file or directory with --base-checkpoint. The loader checks model type, tensor keys, and tensor shapes; it does not require a particular file SHA-256 or training-machine path. The base hash recorded in metadata is training provenance. Load this adapter on top of the base; optimizer state and base weights are not included.
The policy consumes three RGB cameras (top, left wrist, right wrist), a 14-dimensional native state, and the task prompt. Joint and gripper actions remain absolute targets in the dataset's original units, with no Aloha unit conversion or state subtraction. Quantile normalization uses the included statistics. Model vectors are padded to 32 dimensions; inference returns the first 14 dimensions. Actions are produced in chunks of 50 frames. Pi0.5 uses discrete state input and a 200-token limit.
Training snapshot
| Setting | Value |
|---|---|
| Backbone | Pi0.5 |
| Saved optimizer step | 14,000 |
| Planned training steps | 50,000 |
| Dataset | Travor278/piperx-put-cube-in-drawer-20260908-87ep |
| Dataset revision | 58bbbd720f6e78d162b8f4bc7077759d34c5162f |
| Data size | 87 episodes, 54,463 frames, 30 FPS |
| Global batch | 16 (4 GPUs, 4 samples per GPU, accumulation 1) |
| Learning rate | Peak 2.5e-5, 1,000 warmup steps, cosine decay to 2.5e-6 |
| Compute precision | bfloat16 |
| Seed | 42 |
| Generator width | 1,024 |
| Adapted layers | 8 VLM layers and 6 action layers |
| Source commit | 1d2bd070bf2a6a9f066601c09ab36cd489264178 |
| Public base commit at run start | e5433de720523f020447e183861a097b902ba7b9 |
The training data was read through an RGB-only LeRobot v2.1-compatible view of the original v3 dataset. Numeric state/action values and RGB videos were preserved; action windows stop at episode boundaries.
sha256sums.txt records the payload checksums. The step-14000 tag identifies this snapshot.