File size: 5,006 Bytes
5a54af7
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
# 0806 / 0807 speed-ablation checkpoints

Five checkpoints published to
[`Alan0928/VLAwithVariousSpeed`](https://huggingface.co/Alan0928/VLAwithVariousSpeed)
under `checkpoints/<config>/<exp>/30000/`, the same layout the repo already used
for the 0510/0511/0513 generation.

All five share the training recipe, so any difference between them is the thing
being ablated and nothing else:

| | |
|---|---|
| Base weights | `pi05_base` (`pi0_fast_base` for the FAST variant) |
| Speeds | 0.75, 1.0, 1.25, 1.5 (online sliding chunks) |
| Steps | 30,000 |
| Batch | 512 |
| LR | cosine, 1,000 warmup, peak 1e-4, decay to 1e-6 |
| Optimizer | AdamW, grad-norm clip 1.0, EMA 0.999 |
| Action horizon | 30 |
| Asset id | `online_sliding_speed_embed_0p75_1p0_1p25_1p5_pi05` |

Norm stats are the source 1.0x statistics for every speed, so all five ship an
identical `norm_stats.json` β€” see the norm-stats note in `CLAUDE.md` for why
that is deliberate.

## What each one isolates

| Config / exp | Varies | Source branch |
|---|---|---|
| `pi05_libero_speed_embed_text`<br>`0806_..._speed_embed_text_...` | reference condition β€” pi0.5 flow matching, speed as instruction text | `0809_ablation_eval` |
| `pi05_libero_speed_embed_none`<br>`0806_..._nospeed_...` | **no speed conditioning.** Same multi-speed data, but no prompt rewrite and no model-side speed module, so the policy has to average over speeds. Adds no new parameters. | `0809_ablation_eval` |
| `pi05_libero_speed_embed_text_noseg`<br>`0806_..._noseg_text_...` | **no motion segmentation** in the re-timer. Isolates what homogeneous-motion segmentation contributes; speed still arrives as text. | `0809_ablation_eval` |
| `pi05_libero_speed_embed_text_regression`<br>`0806_..._regression_text_...` | **L1 action regression** instead of flow matching (learnable action queries, no timestep, no adaRMS). | `0809_arch_experiments_eval` |
| `pi0_fast_libero_speed_embed_text`<br>`0807_pi0fast_...` | **pi0-FAST** β€” autoregressive over FAST action tokens. `action_dim=32` and `max_token_len=250` are load-bearing (see the config comment). | `0809_arch_experiments_eval` |

## LIBERO results

Overall success rate across all four suites, 500 episodes per suite (2,000 per
speed), 50 trials per task. `steps` is `mean_steps_all`.

| Checkpoint | 0.75x | 1.0x | 1.25x | 1.5x |
|---|---|---|---|---|
| `text_noseg` | 95.9% / 212 steps | 97.6% / 159 | 97.4% / 136 | 96.6% / 127 |
| `none` (no conditioning) | β€” | 96.8% / 152 | β€” | β€” |
| `text_regression` | 92.2% / 224 | 93.2% / 168 | 93.0% / 147 | 90.7% / 141 |
| `pi0_fast` | 91.0% / 225 | 93.5% / 167 | β€” | β€” |
| `text` (reference) | β€” | β€” | β€” | β€” |

**The table is incomplete, and the gaps are missing evals rather than failures:**

- The **reference `text` checkpoint has no eval results at all** in either tree.
  Every comparison above is therefore against a missing baseline β€” the ablations
  cannot yet be read as deltas.
- `none` was only run at 1.0x. That is the one speed where "no conditioning" is
  not obviously handicapped, so it is the least informative point; 0.75/1.25/1.5
  are the ones that would show whether conditioning matters.
- `pi0_fast` is missing 1.25x and 1.5x.

Per-suite numbers, per-episode JSON and run manifests are committed under
`results/ablation_0806/` and `results/ablation_arch/` on the two branches.
Rollout videos are gitignored (778M / 1.1G of mp4) and are not on HF either.

## What is not published

`optimizer.pt` is **not** uploaded for these five, unlike the 0510/0511/0513
checkpoints which do carry theirs. It is ~13GB each and only matters if you
resume training; skipping it took this upload from ~94GB to ~35GB. The files are
still on the workstation if a resume is ever needed.

## Using one

```bash
hf download Alan0928/VLAwithVariousSpeed \
  --include "checkpoints/pi05_libero_speed_embed_text_noseg/*" \
  --local-dir ./hf_ckpts

TORCH_COMPILE_DISABLE=1 CUDA_VISIBLE_DEVICES=0 uv run python scripts/serve_policy.py policy:checkpoint \
  --policy.config pi05_libero_speed_embed_text_noseg \
  --policy.dir ./hf_ckpts/checkpoints/pi05_libero_speed_embed_text_noseg/0806_pi05_online_sliding_noseg_text_0p75_1p0_1p25_1p5/30000 \
  --policy.asset-id online_sliding_speed_embed_0p75_1p0_1p25_1p5_pi05 \
  --port 8000
```

`TORCH_COMPILE_DISABLE=1` is not optional β€” the training configs use
`pytorch_compile_mode='max-autotune'` and leaving compile on stalls server
startup for 5–15 minutes.

Serving the regression and pi0-FAST variants additionally needs the
`load_pytorch` architecture dispatch from the `0809_arch_experiments_eval`
branch; without it the loader always builds `PI0Pytorch` and dies on a
`state_dict` mismatch.

## Reproducing the upload

```bash
HF_TOKEN=... vpn python3 scripts/upload_checkpoints_to_hf.py
```

Idempotent β€” size-matched files are skipped, so a run interrupted by a proxy
drop can just be re-run. `vpn` is needed because huggingface.co is not reachable
directly from the office network.