File size: 2,293 Bytes
fc14206
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
41eb1b3
fc14206
 
 
 
 
 
 
 
 
a6aa576
fc14206
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
---
license: other
library_name: pytorch
tags:
- video-generation
- flow-matching
- distillation
- robotics
- wan
---

# WAM-Policy — Video Forward Flow-map Distillation (v2)

Flow-map–distilled **video student** models for WAM-Policy v2, distilled from a frozen
**LingBot-VA RoboTwin** teacher (WAN 2.2 `WanTransformer3DModel`). The student learns a
two-time **average-velocity flow-map** so it can sample video latents in **4 steps**
instead of the teacher's 25.

## Contents

| Folder | Student | Train steps | Code |
|---|---|---|---|
| `video_ffm_90k/` | 90,000-step student | 90k | new dual-head v2 |
| `video_ffm_50k/` | 50,000-step student | 50k | original video-head FFM |
| `video_ffm_5k/`  | 5,000-step student  | 5k  | new dual-head v2 |

Each folder contains:
- `teacher_25step.mp4` — teacher sampled with **25** steps
- `student_4step.mp4` — student sampled with **4** steps
- `teacher_vs_student.mp4` — side-by-side comparison
- `offline_video_stream_metadata.json` — render metadata (8 latent chunks, fps 10)
- `checkpoint.pt` — full model checkpoint (frozen teacher backbone + trained flow-map adapters, ~12 GB)
- `adapters_only.pt` — all **non-backbone** weights (the trained flow-map conditioners + video/action heads; everything except the frozen `base_model.blocks` and `condition_embedder`), ~267 MB. Reattach onto the public teacher backbone to reconstruct the student.

## Teacher

Distilled from [robbyant/lingbot-va-posttrain-robotwin](https://huggingface.co/robbyant/lingbot-va-posttrain-robotwin)
(WAN 2.2 video transformer + VAE), kept **frozen**. Only `video_time_pair_adapter` and
`video_head_adapter` are trained.

## Data

RoboTwin WAN 2.2 precomputed video latents
([robbyant/robotwin-clean-and-aug-lerobot](https://huggingface.co/datasets/robbyant/robotwin-clean-and-aug-lerobot)),
an 82,414-sample latent manifest.

## Notes

- `checkpoint.pt` bundles the frozen teacher backbone, hence ~12 GB; the *trained* part is
  only the two small adapter modules (`adapters_only.pt`).
- Trained and rendered on an AMD Instinct **MI355X** (ROCm 7.0, `lingbot_attn_mode=torch`).
- The point of the comparison: the **4-step student** vs the **25-step teacher** — flow-map
  distillation trades sampling steps for a single learned two-time transition.