RockTalk
/

Lance-3B-Video-MLX

@@ -3,97 +3,171 @@ license: apache-2.0
 base_model:
 - bytedance-research/Lance
 - Qwen/Qwen2.5-VL-3B-Instruct
-pipeline_tag: any-to-any
 library_name: mlx
 tags:
 - multimodal
 - mlx
 - apple-silicon
 - image-generation
 - video-generation
-- image-editing
-- video-understanding
-- any-to-any
 - port
 ---
 # Lance-3B-Video-MLX
-A native [MLX](https://github.com/ml-explore/mlx) port of [ByteDance's Lance](https://huggingface.co/bytedance-research/Lance) — a 3B-parameter unified multimodal model for image and video generation, editing, and understanding.
-Built on top of the Qwen2.5-VL-3B-Instruct backbone, with Lance's custom multi-task adapters and a Wan 2.2 VAE.
-## Status
-**Weight conversion is complete** — all tensors from the upstream PyTorch
-checkpoint are present in MLX safetensors layout, verified bit-exact on
-sampled tensors. **Inference wrapper is not yet runnable.**
-| Component | Status |
 |---|---|
-| Weight conversion (PT → MLX safetensors, layout + name remaps) | ✅ DONE — bit-exact spot check |
-| `modeling_utils` (TimestepEmbedder, PositionEmbedding3D, MLP, sincos tables) | ✅ DONE |
-| `vae_wan22` — image-mode encode/decode | ✅ DONE |
-| `vae_wan22` — video streaming feat-cache | ⏳ PENDING |
-| `lance.py` adapters | ⚠ PARTIAL — primitives present; patchify path needs source-faithful rewrite |
-| Lance MoE-gen attention (`_moe_gen` weights bundled, 505 tensors) | ⏳ NOT YET WRAPPED |
-| Lance QK-norm extension (`q_norm`/`k_norm` weights bundled, 73 tensors) | ⏳ NOT YET WRAPPED |
-| Flow-matching sampler / CFG | ⏳ STUB |
-| X→T autoregressive / NaViT | ⏳ Phase 2 |
 ## Files
-- `model.safetensors` — Lance 3B **video variant** LLM + adapters (Qwen2.5-VL language model, vae2llm/llm2vae, time_embedder, latent_pos_embed), MLX-layout, **26.4 GB** (1411 tensors, ~7.1B params)
-- `vit.safetensors` — Qwen2.5-VL ViT visual encoder, MLX-layout (NTHWC conv weights), **1.25 GB** (390 tensors, ~668M params, fp16 — bundled here for offline use)
-- `vae.safetensors` — Wan 2.2 VAE, MLX-layout, **2.62 GB** (196 tensors, ~705M params). Converted from the upstream `Wan2.2_VAE.pth` pickle.
-- `config.json` — distilled architecture config + embedded Qwen2.5-VL sub-config
-- `vit_config.json` — Qwen2.5-VL ViT sub-config
-- `tokenizer.json`, `tokenizer_config.json`, `vocab.json`, `merges.txt`, `generation_config.json` — copied verbatim from upstream
-## Hardware
-Targets Apple Silicon with unified memory. Verified on M3 Ultra (512 GB). Lower-RAM Macs may need to run the LLM forward only (no joint backbone + VAE).
-## Loading
 ```python
 import mlx.core as mx
-import mlx.nn as nn
 from lance_mlx.lance import Lance, LanceConfig
-from mlx_vlm.models.qwen2_5_vl.config import ModelConfig as Qwen25VLConfig
-import json
-from lance_mlx import load_lance
-model, cfg = load_lance("./")
-# cfg['_loaded_into_model'] reports how many tensors landed in the wrapper.
-# cfg['_vae_weights'] holds the Wan VAE keys (load into a separate vae module).
-# cfg['_moe_gen_weights'] holds Lance's generation-path weights, parked until
-# a MoE-aware wrapper is available.
 ```
-## Citation
-```bibtex
-@article{lance2026,
-  title   = {Lance: Unified Multimodal Modeling by Multi-Task Synergy},
-  author  = {Fu, Fengyi and Huang, Mengqi and Wu, Shaojin and Jiang, Yunsheng and Huo, Yufei and Guo, Jianzhu and others},
-  journal = {arXiv preprint arXiv:2605.18678},
-  year    = {2026},
-  url     = {http://arxiv.org/abs/2605.18678}
-}
 ```
 ## License
-Apache-2.0, inherited from upstream `bytedance-research/Lance`.
-## Acknowledgments
-- ByteDance Research for the original Lance training and PyTorch release
-- The `mlx` and `mlx-vlm` teams at Apple
-- Qwen team for Qwen2.5-VL-3B-Instruct
----
-**Port status reporting honestly:** this repo currently provides MLX-format weights with verified-loading scaffolding. Inference sampling (T2I/T2V) is a follow-up release; the building blocks are in place but the diffusion loop has not been parity-validated end-to-end yet. Pull requests welcome.

 base_model:
 - bytedance-research/Lance
 - Qwen/Qwen2.5-VL-3B-Instruct
+pipeline_tag: text-to-video
 library_name: mlx
 tags:
 - multimodal
 - mlx
 - apple-silicon
+- text-to-image
 - image-generation
 - video-generation
+- diffusion
+- flow-matching
+- moe
+- qwen2_5_vl
+- wan
 - port
 ---
 # Lance-3B-Video-MLX
+Video variant of [Lance-3B-MLX](https://huggingface.co/RockTalk/Lance-3B-MLX). First native [MLX](https://github.com/ml-explore/mlx) port of [ByteDance Research's Lance](https://huggingface.co/bytedance-research/Lance) — a 3B-parameter unified multimodal model for image/video generation, editing, and understanding. Runs natively on Apple Silicon, no CUDA required.
+The architecture is **Qwen2.5-VL-3B + parallel MoE-gen experts + Wan 2.2 VAE**. Lance uses a "Mixture-of-Tokens" routing: every attention block and MLP has a parallel `*_moe_gen` branch. Text tokens go through normal weights; VAE-latent (generation) tokens go through the `_moe_gen` weights, in the same forward pass.
+## What works
+| Capability | Status |
+|---|---|
+| Text-to-image (T2I), single image, CFG | ✅ Working, verified |
+| Strict load of all 1021 LLM/adapter tensors | ✅ Working |
+| Wan 2.2 VAE encode/decode (T=1) | ✅ Working (uses [RockTalk/Wan2.2-VAE-MLX](https://huggingface.co/RockTalk/Wan2.2-VAE-MLX)) |
+| Flow-matching denoising loop | ✅ Working |
+| Classifier-free guidance | ✅ Working |
+| 3D mrope position embeddings | ✅ Working |
+| MoE-gen routing (per-token attention + MLP + layernorm) | ✅ Working |
+| Text-to-video (T2V) | ⚠ Weights ready (31 latent frames × 64² positional grid); needs Wan 2.2 VAE T>1 streaming cache to materialize video frames end-to-end |
+| Image/video editing (TI2I, TIV2V) | ⏳ Phase 2 — needs ViT integration |
+| X→T (image/video understanding) | ⏳ Phase 2 — needs AR sampling loop + KV cache |
+## Sample generations
+Verified on M4 Studio (128 GB). 30 steps, CFG=4, 512×512:
+| Prompt | Output |
 |---|---|
+| *"a photo of a sunset over mountains"* | ![sunset](samples/sunset_mountains.png) |
+| *"a fluffy orange cat sitting on a wooden chair, photorealistic"* | ![cat](samples/orange_cat_chair.png) |
+| *"a majestic snowy mountain peak with a dramatic blue sky and clouds"* | ![mountain](samples/snowy_peak.png) |
+## Performance
+Measured on M4 Studio (128 GB) at CFG=4 (one conditional + one unconditional forward per step):
+| Resolution | Steps | Per-step | Total sample | VAE decode |
+|---|---|---|---|---|
+| 256×256 | 24 | ~400 ms | ~9.6 s | ~0.1 s |
+| 512×512 | 30 | ~1.2 s  | ~36 s | ~0.5 s |
+First-call kernel-compile penalty: ~few seconds per new resolution.
+## Differences vs Lance-3B-MLX
+This is the same architecture as the image variant, with two differences:
+- `model.safetensors`: 26.5 GB (vs 23 GB) — extra weights for multi-frame attention
+- `latent_pos_embed.pos_embed`: 31 × 64 × 64 = 126,976 positions (vs 1 × 64 × 64 = 4,096) — supports up to 31 latent frames (≈ 121 video frames @ 4× temporal downsample)
+T2I via this checkpoint works the same as Lance-3B-MLX. T2V will work once the Wan 2.2 VAE temporal streaming cache is implemented (v0.1.0 of [RockTalk/Wan2.2-VAE-MLX](https://huggingface.co/RockTalk/Wan2.2-VAE-MLX)).
 ## Files
+| File | Size | Description |
+|---|---|---|
+| `model.safetensors` | 26.5 GB | LLM (Qwen2.5-VL with MoE-gen) + Lance adapters, 1021 tensors |
+| `vit.safetensors` | 1.25 GB | Qwen2.5-VL ViT (for understanding mode — Phase 2) |
+| `vae.safetensors` | 2.62 GB | Wan 2.2 VAE (older keying — for compatibility; the standalone [RockTalk/Wan2.2-VAE-MLX](https://huggingface.co/RockTalk/Wan2.2-VAE-MLX) uses cleaner keys and is recommended) |
+| `config.json` | — | Distilled architecture config |
+| `tokenizer.json`, `vocab.json`, `merges.txt` | — | Qwen2.5-VL tokenizer, verbatim |
+| `samples/*.png` | — | Verified T2I outputs from this checkpoint |
+## Usage
+Requires `mlx >= 0.29`, `mlx-vlm >= 0.3`, `numpy`, `einops`, `transformers`, `pillow`, and the [`lance-mlx`](https://github.com/RockTalk/Lance-MLX) companion repo for the `Lance` Python class.
+```bash
+pip install mlx mlx-vlm numpy einops transformers pillow
+```
 ```python
 import mlx.core as mx
 from lance_mlx.lance import Lance, LanceConfig
+from lance_mlx.vae_wan22 import Wan2_2_VAE
+# Build + strict-load (see tools/lance_t2i.py in the companion repo for the
+# full builder; LanceConfig takes a Qwen2.5-VL ModelConfig built from
+# config.json).
+model = Lance(lance_cfg)
+model.load_weights(list(mx.load("model.safetensors").items()), strict=True)
+vae = Wan2_2_VAE(z_dim=48, c_dim=160, dim_mult=(1, 2, 4, 4),
+                 temperal_downsample=(False, True, True))
+vae.model.load_weights(list(mx.load("vae.safetensors").items()), strict=True)
+# Sample
+latent = model.sample_t2i(
+    prompt_token_ids=text_ids,                  # (P,) int32 from tokenizer (no specials)
+    latent_shape=(1, 32, 32),                   # (T_lat, H_lat, W_lat) for 512×512 image
+    special_token_ids={"bos": 151644, "eos": 151645,
+                       "start_of_image": 151652, "end_of_image": 151653,
+                       "image_token_id": 151655},
+    num_steps=30, timestep_shift=3.5, cfg_scale=4.0, seed=0,
+)
+img = vae.decode(latent)                        # (1, 1, 512, 512, 3) in [-1, 1]
+```
+End-to-end script: `tools/lance_t2i.py` in the [companion repo](https://github.com/RockTalk/Lance-MLX).
+## How the MoE-gen routing is implemented in MLX
+Lance's checkpoint contains *two* sets of weights per Qwen2 block:
+```
+self_attn.{q,k,v,o}_proj         self_attn.{q,k,v,o}_proj_moe_gen
+self_attn.{q,k}_norm             self_attn.{q,k}_norm_moe_gen
+mlp.{gate,down,up}_proj          mlp_moe_gen.{gate,down,up}_proj
+input_layernorm                  input_layernorm_moe_gen
+post_attention_layernorm         post_attention_layernorm_moe_gen
 ```
+For T2I/T2V the sequence layout is:
 ```
+<|im_start|> [prompt tokens] <|im_end|> <|vision_start|> [N latent placeholders] <|vision_end|>
+                                                          └──── routed through moe_gen ────┘
+                                                          ↑ everything else: normal weights
+```
+The MLX port (`qwen2_navit_mlx.py`) routes by slicing the sequence into the latent slab vs the surrounding text, applying the appropriate expert to each slab, and concatenating. mrope position ids continue to flow normally across both slabs (with axis-T/H/W coordinates only varying inside the latent slab).
+## Conversion source
+Converted from `bytedance-research/Lance/Lance_3B/*` using the open-source pipeline at https://github.com/RockTalk/Lance-MLX (`tools/convert_weights.py`). Layout transforms:
+- Conv weights: PT `(O, I, [T,] H, W)` → MLX `(O, [T,] H, W, I)`
+- Embedding weights: shape preserved
+- `lm_head.weight` tied to `embed_tokens.weight` (Qwen default)
+- All `*_moe_gen.*` keys copied verbatim under the same names
 ## License
+Apache 2.0, inherited from upstream `bytedance-research/Lance`. The Wan 2.2 VAE component is also Apache 2.0 from Alibaba's Wan team.
+## Acknowledgements
+- **ByteDance Research** — original Lance training + PT release
+- **Qwen team** — Qwen2.5-VL-3B-Instruct backbone
+- **Alibaba Wan team** — Wan 2.2 VAE training
+- **Apple `mlx` and `mlx-vlm` teams** — the underlying frameworks
+- **This MLX port** — RockTalk
+## Citation
+```bibtex
+@misc{lance_mlx,
+  title  = {Lance-3B-MLX — First MLX port of ByteDance's Lance},
+  author = {RockTalk},
+  year   = {2026},
+  url    = {https://huggingface.co/RockTalk/Lance-3B-MLX}
+}
+```

config.json CHANGED Viewed

@@ -67,14 +67,15 @@
   },
   "latent_patch_size": [
     1,
-    2,
-    2
   ],
-  "max_latent_size": 32,
-  "max_num_frames": 25,
   "latent_channel": 48,
   "vae_downsample_spatial": 16,
   "vae_downsample_temporal": 4,
   "connector_act": "gelu_pytorch_tanh",
-  "timestep_shift": 3.5
 }

   },
   "latent_patch_size": [
     1,
+    1,
+    1
   ],
+  "max_latent_size": 64,
+  "max_num_frames": 120,
   "latent_channel": 48,
   "vae_downsample_spatial": 16,
   "vae_downsample_temporal": 4,
   "connector_act": "gelu_pytorch_tanh",
+  "timestep_shift": 3.5,
+  "max_num_latent_frames": 31
 }