pi0.5-s8bNBi2ZQghb — OpenRoboto AXIS v1.0 · π0.5 (competition 6)
Full openpi (JAX) checkpoint for the AXIS v1.0 simulation season (Franka Panda, 30 MuJoCo tasks,
evaluator config pi05_axis_joint, norm stats at assets/axis-v0.1-task501-runtime-v1/norm_stats.json).
Parent (disclosed)
- Parent:
pfenzi/pi0.5-DvM7hJjfDcyw@c0e6cbb795aa2ad3c302f1e519efb9e581a36198— the AXIS competition-6 champion at training time (0.9733). - 47 of 51 tensors are byte-identical to the parent; normalization statistics are byte-identical. Architecture, tokenizer, action horizon and the evaluator's sampler are unchanged.
Method — per-task first-step offset, time-localized
The parent's first Euler step of the 10-step flow sampler is (near-)independent of the input noise (a first-step "noise collapse"), so each task's trajectory is determined by the first-step point x₀.₉(obs).
- For the parent's weak tasks we searched, with the parent itself and the official evaluator on our own policy seeds, for a small per-task offset of that point, x₀.₉ ← x₀.₉ + σ·N(seed), keeping the parent's velocity field for t ≤ 0.9. Offsets were accepted only if they succeeded under small action perturbations (robustness check): task 22 (σ=0.5), task 501 (σ=0.5), task 42 (σ=0.1). All other tasks keep the parent's point.
- On-policy rollouts of the parent with those offsets (tasks 22/501/42) and without (all other tasks), with small random action perturbations for state coverage. No evaluation trial addresses were used.
- Objective: the first Euler step from any noise must land on the target point (parent point + task offset,
or the unchanged parent point for the other tasks). Only the adaRMS time-conditioning kernels
(
pre_attention_norm_1,pre_ffw_norm_1,final_norm_1modulation Dense kernels andtime_mlp_out) are trained; every update is projected onto the orthogonal complement of the conditioning vectors at t = 0.9 … 0.1, so the parent's velocity field for t ≤ 0.9 is preserved exactly (up to rounding). 3,000 AdamW updates, batch 32, LR 1e-4, JAX/openpi.
Local results (official evaluator, 30 tasks × 20 trials, our policy seeds)
| model | 777 | 888 | 999 | 1111 |
|---|---|---|---|---|
| this model | 591 | 596 | 594 | 593 |
| parent pfenzi/pi0.5-DvM7hJjfDcyw | 581 | 580 | 574 | 589 |
Main change: task 22 (Grab Can) 9/20 and 7/20 → 20/20 on seeds 777/888.
License
Weights derive from openpi π0.5 (Apache-2.0) and PaliGemma (Gemma Terms of Use).
Model tree for deepmaster/pi0.5-s8bNBi2ZQghb
Base model
pfenzi/pi0.5-DvM7hJjfDcyw