File size: 3,343 Bytes
b7fe454 0f4d075 b7fe454 8015ff5 b7fe454 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 | ---
license: mit
tags:
- robotics
- franka
- manipulation
- video-diffusion
- fastwam
---
# deploy_public β FastWAM / FlashWAM lift_new policy weights (Franka "pick up the red cube")
Generation-2 weights (2026-07-12) consumed by the
[`sleepmastergx/deploy`](https://github.com/sleepmastergx/deploy) deployment
code (`fetch_weights.py` downloads gen-2 from here; gen-1 and the shared base
pieces β VAE, ActionDiT, T5 text-embed cache β live in `SleepMastger/deploy`).
Layout: `runs/<run>/<timestamp>/checkpoints/weights/step_*.pt`, each run with
**its own** `dataset_stats.json` (min/max normalization) and resolved training
`config.yaml`. β οΈ Always pair a checkpoint with the `dataset_stats.json` from
the SAME run β foreign stats mis-scale actions.
## The four runs β deploy with `deploy_v2_new/`
Trained on [SleepMastger/lift_new](https://huggingface.co/datasets/SleepMastger/lift_new)
(103 episodes @ **10 Hz**, REAL gripper widths). Contract: execute xyz deltas
at 10 Hz; proprio `state[6:8] = [w/2, -w/2]` (meters). 30 epochs, global
batch 32, lr 1e-4 cosine. Five checkpoints per run, **named by true epoch**
(saved every 500 steps at 265 steps/epoch, so most land mid-epoch):
`epoch_05.7.pt` / `epoch_09.4.pt` / `epoch_15.1.pt` / `epoch_20.8.pt` /
`epoch_30.0.pt`. (The last three also exist under their original step names
`step_002500/005500/007950`.) Training loss β robot success in this
103-episode overfitting regime, so A/B the earlier epochs too.
| Run | Architecture | Init | Final action loss |
|---|---|---|---|
| `runs/lift_new_flashwam_scratch/` | FlashWAM M1_FusedKV_RopeFixed (1-layer action expert, fused_kv, fixed RoPE) | Wan2.2-TI2V-5B base + random action expert | 0.0155 |
| `runs/lift_new_flashwam_ft/` | same | LIBERO M1_FusedKV_RopeFixed step_021700 (format-converted) | 0.0061 |
| `runs/lift_new_fastwam_scratch/` | Original FastWAM (30/30 MoT, 6.0B) | Wan2.2 base + ActionDiT interp init | 0.0061 |
| `runs/lift_new_fastwam_ft/` | same | LIBERO release checkpoint | 0.0055 |
FlashWAM (fused_kv) checkpoints require fastwam code at/after commit
`8d9e040` (2026-07-08 fused_kv reparameterization) β the vendored copy in the
GitHub repo qualifies.
## `video_expert/` β Wan video DiT before vs after LIBERO training
The 5.0B Wan2.2 video expert in three states (see `video_expert/README.md`
for format details):
- `wan22_base_before_libero/` β pristine Wan2.2-TI2V-5B DiT (before any robot
training; fp32 safetensors, verbatim from `Wan-AI/Wan2.2-TI2V-5B`)
- `wan_video_expert_after_libero_fastwam_release.pt` β after LIBERO training,
original FastWAM release lineage (bf16, extracted from the MoT checkpoint)
- `wan_video_expert_after_libero_flashwam_fusedkv.pt` β after LIBERO training,
FlashWAM M1_FusedKV_RopeFixed lineage (bf16)
## Notes
- The ~18.6 GB Wan2.2-TI2V-5B base DiT is NOT re-hosted β fetched from the
public `Wan-AI/Wan2.2-TI2V-5B`. VAE / ActionDiT / T5 cache: in
`SleepMastger/deploy` (`fetch_weights.py` assembles everything).
- All policies: two 256Β² RGB cameras (agentview + wrist), proprio [8], action
chunk [32, 7] β xyz delta meters, rpy β 0 (cannot rotate), gripper Β±1 after
the server-side chain (β1=open, +1=close).
- Training curves: wandb `huaweiwam/fastwam-realrobot` (runs `fpl22etl`,
`xnndko8y`, `1ju16696`, `k42n0ba4`).
|