pi0.5-s8bNBi2ZQghb — OpenRoboto AXIS v1.0 · π0.5 (competition 6)

Full openpi (JAX) checkpoint for the AXIS v1.0 simulation season (Franka Panda, 30 MuJoCo tasks, evaluator config pi05_axis_joint, norm stats at assets/axis-v0.1-task501-runtime-v1/norm_stats.json).

Parent (disclosed)

  • Parent: pfenzi/pi0.5-DvM7hJjfDcyw @ c0e6cbb795aa2ad3c302f1e519efb9e581a36198 — the AXIS competition-6 champion at training time (0.9733).
  • 47 of 51 tensors are byte-identical to the parent; normalization statistics are byte-identical. Architecture, tokenizer, action horizon and the evaluator's sampler are unchanged.

Method — per-task first-step offset, time-localized

The parent's first Euler step of the 10-step flow sampler is (near-)independent of the input noise (a first-step "noise collapse"), so each task's trajectory is determined by the first-step point x₀.₉(obs).

  1. For the parent's weak tasks we searched, with the parent itself and the official evaluator on our own policy seeds, for a small per-task offset of that point, x₀.₉ ← x₀.₉ + σ·N(seed), keeping the parent's velocity field for t ≤ 0.9. Offsets were accepted only if they succeeded under small action perturbations (robustness check): task 22 (σ=0.5), task 501 (σ=0.5), task 42 (σ=0.1). All other tasks keep the parent's point.
  2. On-policy rollouts of the parent with those offsets (tasks 22/501/42) and without (all other tasks), with small random action perturbations for state coverage. No evaluation trial addresses were used.
  3. Objective: the first Euler step from any noise must land on the target point (parent point + task offset, or the unchanged parent point for the other tasks). Only the adaRMS time-conditioning kernels (pre_attention_norm_1, pre_ffw_norm_1, final_norm_1 modulation Dense kernels and time_mlp_out) are trained; every update is projected onto the orthogonal complement of the conditioning vectors at t = 0.9 … 0.1, so the parent's velocity field for t ≤ 0.9 is preserved exactly (up to rounding). 3,000 AdamW updates, batch 32, LR 1e-4, JAX/openpi.

Local results (official evaluator, 30 tasks × 20 trials, our policy seeds)

model 777 888 999 1111
this model 591 596 594 593
parent pfenzi/pi0.5-DvM7hJjfDcyw 581 580 574 589

Main change: task 22 (Grab Can) 9/20 and 7/20 → 20/20 on seeds 777/888.

License

Weights derive from openpi π0.5 (Apache-2.0) and PaliGemma (Gemma Terms of Use).

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Model tree for deepmaster/pi0.5-s8bNBi2ZQghb

Finetuned
(1)
this model