Miowtion

Veda-MiniMax-H3 R2VA (Preview)

Demo · Project Page · Code · Paper · Deployment Guide

Introduction

Veda is a learned sparse-attention method for video diffusion models. This checkpoint brings Veda to MiniMax-H3 reference-to-audio-video generation. It selects visual reference and generated video tiles independently; text, VLM conditioning, and audio attention remain dense. The predictor is not a LoRA and leaves the backbone weights unchanged.

This preview was fine-tuned for 600 updates from the T2VA predictor on 5.17-second R2VA clips. It uses the MiniMax-H3 Turbo LoRA for eight-step generation, with up to 32 visual-reference tiles and 32 current-video tiles per query tile and head. The predictor is stored in FP8 and dequantized to BF16 when loaded.

Highlights

  • Learned sparse attention. Visual reference and generated-video attention are selected independently; text, VLM conditioning, and audio stay dense.
  • Fixed visual budgets. The released preview uses 32 reference tiles and 32 current-video tiles per query tile and head.
  • Plug-and-play predictor. The FP8 predictor is separate from the H3 backbone and Turbo LoRA weights.

Samples

The examples below use the same seed, prompt, geometry, and Turbo LoRA for Dense and Veda. Each generated output is 1344×768, 124 frames, and about 5.17 seconds. Dense (left) and Veda (right) are stitched into one synchronized comparison video, with attention and denoising speedups in its top title bar. Each comparison uses the original Dense audio track.

TaskReferencesDense (left) · Veda (right)
Pink suit and lamb
A man in a bright pink suit holds a black lamb in a sunlit pasture, speaks a short line, and keeps the source soundtrack.
Video 1 · appearance and scene

Audio 1 · source soundtrack

Audio 2 · male voice reference

Open comparison · Open Dense video · Open Veda video
Anime train by the coast
An anime traveler watches a bright coastline pass outside a moving train window.
Picture 1 · character and coastal scene
Anime traveler beside a coastal train window
open reference

Open comparison · Open Dense video · Open Veda video
Anime twilight carnival
An anime woman walks through a fairground as the Ferris wheel and lights turn on at blue hour.
Picture 1 · character
Anime character reference
open character reference
Picture 2 · fairground
Anime fairground reference
open fairground reference

Open comparison · Open Dense video · Open Veda video
Aurora over the lake
A green-violet aurora slowly grows above a mountain lake and appears in the calm reflection.
Picture 1 · lake and sky
Aurora over a mountain lake
open lake reference
Video 1 · aurora motion

Open comparison · Open Dense video · Open Veda video
Red fox in autumn woods
A red fox crosses a warm autumn woodland, slows, and looks toward the camera.
Picture 1 · fox appearance
Red fox appearance reference
open fox reference
Video 1 · fox motion

Open comparison · Open Dense video · Open Veda video
Perfume by the pool
A glass perfume bottle catches warm light beside rippling water in a stone courtyard.
Picture 1 · product and pool
Perfume bottle beside a pool
open product reference

Open comparison · Open Dense video · Open Veda video

Performance

Measured with eight-step Turbo LoRA inference on one RTX PRO 6000 Blackwell GPU. Veda uses fixed budgets of 32 reference and 32 current-video tiles.

GPU Clip Attention speedup End-to-end speedup
RTX PRO 6000 Blackwell 16:9 · 5.17 s 2.42× 1.69×
RTX PRO 6000 Blackwell 9:16 · 5.17 s 2.65× 1.67×
RTX PRO 6000 Blackwell 1:1 · 5.17 s 2.06× 1.37×
RTX PRO 6000 Blackwell 4:3 · 5.17 s 2.19× 1.54×
RTX PRO 6000 Blackwell 16:9 · 10.1 s 2.84× 2.05×
RTX PRO 6000 Blackwell 16:9 · 14.4 s 2.90× 2.20×

Usage

See AGENTS.md for installation and GPU-specific deployment notes. The commands below use the R2VA code on the main branch.

git clone --recurse-submodules --branch main https://github.com/veda-sparse/Miowtion.git
cd Miowtion
pip install -e '.[gpu,encode]'

hf download MiniMaxAI/MiniMax-H3 --include "model_index.json" "Ref2VA/*" \
  --local-dir weights/MiniMax-H3
hf download larryvrh/MiniMax-H3-Turbo-Lora \
  minimax_h3_turbo_v4_step600_ema.safetensors --local-dir weights/turbo_lora
hf download Veda-Sparse/Minimax-H3-R2VA-Veda-Preview \
  --local-dir weights/veda/h3-r2va-preview

Prepare the reference files, then encode and generate:

python scripts/prepare_official_r2va_demo.py \
  --out artifacts/examples/minimax_h3_ref2va
python scripts/encode_samples.py --root weights/MiniMax-H3 \
  --manifest artifacts/examples/minimax_h3_ref2va/manifests/prompts.jsonl \
  --out artifacts/samples/demo
python scripts/generate.py --config configs/infer_r2va_preview.yaml \
  --predictor weights/veda/h3-r2va-preview/minimax_h3_r2va_veda_preview_fp8.safetensors \
  --attention dense veda

License

This model inherits the MiniMax H3 Community License from its base model.

Citation

@inproceedings{han2026veda,
  title={Veda: Scalable Video Diffusion via Distilled Sparse Attention},
  author={Han, Shihao and Yang, Hao and Hu, Xinting and Mei, Xiaofeng
          and Jiang, Yi and Qi, Xiaojuan},
  booktitle={International Conference on Machine Learning (ICML)},
  year={2026}
}
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Veda-Sparse/Minimax-H3-R2VA-Veda-Preview

Finetuned
(163)
this model

Paper for Veda-Sparse/Minimax-H3-R2VA-Veda-Preview