Veda-MiniMax-H3 R2VA (Preview)
Demo · Project Page · Code · Paper · Deployment Guide
Introduction
Veda is a learned sparse-attention method for video diffusion models. This checkpoint brings Veda to MiniMax-H3 reference-to-audio-video generation. It selects visual reference and generated video tiles independently; text, VLM conditioning, and audio attention remain dense. The predictor is not a LoRA and leaves the backbone weights unchanged.
This preview was fine-tuned for 600 updates from the T2VA predictor on 5.17-second R2VA clips. It uses the MiniMax-H3 Turbo LoRA for eight-step generation, with up to 32 visual-reference tiles and 32 current-video tiles per query tile and head. The predictor is stored in FP8 and dequantized to BF16 when loaded.
Highlights
- Learned sparse attention. Visual reference and generated-video attention are selected independently; text, VLM conditioning, and audio stay dense.
- Fixed visual budgets. The released preview uses 32 reference tiles and 32 current-video tiles per query tile and head.
- Plug-and-play predictor. The FP8 predictor is separate from the H3 backbone and Turbo LoRA weights.
Samples
The examples below use the same seed, prompt, geometry, and Turbo LoRA for Dense and Veda. Each generated output is 1344×768, 124 frames, and about 5.17 seconds. Dense (left) and Veda (right) are stitched into one synchronized comparison video, with attention and denoising speedups in its top title bar. Each comparison uses the original Dense audio track.
| Task | References | Dense (left) · Veda (right) |
|---|---|---|
| Pink suit and lamb A man in a bright pink suit holds a black lamb in a sunlit pasture, speaks a short line, and keeps the source soundtrack. |
Video 1 · appearance and scene Audio 1 · source soundtrack Audio 2 · male voice reference |
Open comparison · Open Dense video · Open Veda video |
| Anime train by the coast An anime traveler watches a bright coastline pass outside a moving train window. |
Picture 1 · character and coastal scene![]() open reference |
Open comparison · Open Dense video · Open Veda video |
| Anime twilight carnival An anime woman walks through a fairground as the Ferris wheel and lights turn on at blue hour. |
Picture 1 · character![]() open character reference Picture 2 · fairground ![]() open fairground reference |
Open comparison · Open Dense video · Open Veda video |
| Aurora over the lake A green-violet aurora slowly grows above a mountain lake and appears in the calm reflection. |
Picture 1 · lake and sky![]() open lake reference Video 1 · aurora motion |
Open comparison · Open Dense video · Open Veda video |
| Red fox in autumn woods A red fox crosses a warm autumn woodland, slows, and looks toward the camera. |
Picture 1 · fox appearance![]() open fox reference Video 1 · fox motion |
Open comparison · Open Dense video · Open Veda video |
| Perfume by the pool A glass perfume bottle catches warm light beside rippling water in a stone courtyard. |
Picture 1 · product and pool![]() open product reference |
Open comparison · Open Dense video · Open Veda video |
Performance
Measured with eight-step Turbo LoRA inference on one RTX PRO 6000 Blackwell GPU. Veda uses fixed budgets of 32 reference and 32 current-video tiles.
| GPU | Clip | Attention speedup | End-to-end speedup |
|---|---|---|---|
| RTX PRO 6000 Blackwell | 16:9 · 5.17 s | 2.42× | 1.69× |
| RTX PRO 6000 Blackwell | 9:16 · 5.17 s | 2.65× | 1.67× |
| RTX PRO 6000 Blackwell | 1:1 · 5.17 s | 2.06× | 1.37× |
| RTX PRO 6000 Blackwell | 4:3 · 5.17 s | 2.19× | 1.54× |
| RTX PRO 6000 Blackwell | 16:9 · 10.1 s | 2.84× | 2.05× |
| RTX PRO 6000 Blackwell | 16:9 · 14.4 s | 2.90× | 2.20× |
Usage
See AGENTS.md for installation and GPU-specific deployment notes.
The commands below use the R2VA code on the
main branch.
git clone --recurse-submodules --branch main https://github.com/veda-sparse/Miowtion.git
cd Miowtion
pip install -e '.[gpu,encode]'
hf download MiniMaxAI/MiniMax-H3 --include "model_index.json" "Ref2VA/*" \
--local-dir weights/MiniMax-H3
hf download larryvrh/MiniMax-H3-Turbo-Lora \
minimax_h3_turbo_v4_step600_ema.safetensors --local-dir weights/turbo_lora
hf download Veda-Sparse/Minimax-H3-R2VA-Veda-Preview \
--local-dir weights/veda/h3-r2va-preview
Prepare the reference files, then encode and generate:
python scripts/prepare_official_r2va_demo.py \
--out artifacts/examples/minimax_h3_ref2va
python scripts/encode_samples.py --root weights/MiniMax-H3 \
--manifest artifacts/examples/minimax_h3_ref2va/manifests/prompts.jsonl \
--out artifacts/samples/demo
python scripts/generate.py --config configs/infer_r2va_preview.yaml \
--predictor weights/veda/h3-r2va-preview/minimax_h3_r2va_veda_preview_fp8.safetensors \
--attention dense veda
License
This model inherits the MiniMax H3 Community License from its base model.
Citation
@inproceedings{han2026veda,
title={Veda: Scalable Video Diffusion via Distilled Sparse Attention},
author={Han, Shihao and Yang, Hao and Hu, Xinting and Mei, Xiaofeng
and Jiang, Yi and Qi, Xiaojuan},
booktitle={International Conference on Machine Learning (ICML)},
year={2026}
}
- Downloads last month
- -
Model tree for Veda-Sparse/Minimax-H3-R2VA-Veda-Preview
Base model
MiniMaxAI/MiniMax-H3




