SleepMastger's picture
Add files using upload-large-folder tool
b284320 verified
|
Raw
History Blame Contribute Delete
1.78 kB

video_expert — Wan2.2-TI2V-5B video DiT, before vs after LIBERO training

The 5.0B-parameter Wan video expert (30-layer DiT) in three states:

File / dir State Source
wan22_base_before_libero/ BEFORE any robot training — pristine Wan2.2-TI2V-5B video DiT (3 safetensors shards, fp32, verbatim copy of Wan-AI/Wan2.2-TI2V-5B diffusion_pytorch_model-*.safetensors) Wan-AI release
wan_video_expert_after_libero_fastwam_release.pt AFTER LIBERO training — original FastWAM (30/30 MoT) extracted from fastwam_release/libero_uncond_2cam224.pt, step 21700
wan_video_expert_after_libero_flashwam_fusedkv.pt AFTER LIBERO training — FlashWAM M1_FusedKV_RopeFixed run extracted from runs/libero/M1_FusedKV_RopeFixed/2026-07-03_16-13-56/.../step_021700.pt

Two "after" variants exist because each LIBERO run trains its own copy of the video expert (freeze_video_backbone: false); both deviate from the base (e.g. mean|Δ| ≈ 0.0037 on patch_embedding.weight, verified at extraction).

Format of the extracted .pt files

payload = torch.load("wan_video_expert_after_libero_fastwam_release.pt",
                     map_location="cpu", weights_only=True)
sd = payload["state_dict"]   # 825 tensors, bf16, 5.00B params
# keys are the video DiT's own names — patch_embedding.*, blocks.{0..29}.*,
# head.* — i.e. the "mixtures.video." prefix of the FastWAM MoT checkpoint
# stripped. Same naming convention as the base safetensors shards, so
# before/after tensors compare directly by key.
payload["source_checkpoint"], payload["source_step"], payload["description"]

Note dtypes differ across states (base fp32, extracted bf16 as trained) — cast to a common dtype before numeric comparison.