How to use from the
Use from the
Diffusers library
pip install -U diffusers transformers accelerate
import torch
from diffusers import DiffusionPipeline
from diffusers.utils import load_image, export_to_video

# switch to "mps" for apple devices
pipe = DiffusionPipeline.from_pretrained("FreeVideoX/Prism-FreeVideo-Preview", dtype=torch.bfloat16, device_map="cuda")
pipe.to("cuda")

prompt = "A man with short gray hair plays a red electric guitar."
image = load_image(
    "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/diffusers/guitar-man.png"
)

output = pipe(image=image, prompt=prompt).frames[0]
export_to_video(output, "output.mp4")

Prism for FreeVideo (preview)

English · 中文

Prepared weights of Prism (Tencent, MIT) for the Prism preview in FreeVideo. Prism animates a first frame into a 1280 × 720, 8.5-second video with matching sound (speech, sound effects and music). These files are laid out for FreeVideo's block-streaming engine, which runs Prism on a single NVIDIA RTX 30, 40 or 50 series GPU with 12 GiB of VRAM or more. FreeVideo downloads them for you when you select Prism during setup; you do not need to fetch them by hand.

Contents

Folder Size Used by
int8/ 46.6 GiB Every quality level. W8A8 INT8 Prism weights (128-wide Hadamard rotation, quantized from the original weights), the LightX2V Wan2.2-I2V 260412 distillation as an unmerged rank-256 LoRA for the Light level, the UMT5-XXL text encoder (bf16), the Wan2.1 VAE and the DAC audio decoder.
bf16/ 60.8 GiB Optional, the Max level only: the original bf16 Prism weights. It reads the text encoder, tokenizer and VAEs from int8/.

Each folder has a manifest.json with the size and SHA-256 of every file. Files are split per transformer block (expert_high/blocks/NN.safetensors, expert_low/…, audio/…, bridge/…) so that blocks that do not fit in VRAM can stream from RAM or disk.

Source: FrancisRing/Prism revision 347659c562dcc392c45dfe1673c051e7c62f57ac (preview_alpha), MOVA-360p, and the LightX2V 260412 distillation (rank-256 extraction by Kijai).

Quality levels in FreeVideo

Level Recipe Time on one H200
Light 8 distilled steps; the audio follows the original model's features about 6 min
Medium 20 steps, original weights, CFG 5 15.5 min
High 30 steps, original weights, CFG 5 23 min
Max Prism's official 50-step recipe, bf16 weights and exact attention 62.5 min

See the FreeVideo Prism guide for the accelerations, the measured quality against the official sampler, and times under smaller VRAM and RAM.

Licenses

Prism weights and code: MIT, Copyright (C) 2026 Tencent. MOVA, Wan2.1 / Wan2.2 (including the UMT5-XXL text encoder and the Wan VAE) and the LightX2V distillation: Apache-2.0. Descript Audio Codec: MIT. These prepared weights are a derivative of those works and keep their licenses; see LICENSE.

中文

Prism(腾讯,MIT 许可证)为 FreeVideo Prism 预览版准备的权重。Prism 可以从一张首帧生成 1280 × 720、8.5 秒、带同步声音(人声、音效、音乐)的视频。文件按 FreeVideo 的分块流式加载引擎排列,可在 12 GiB 显存及以上的 NVIDIA RTX 30/40/50 系单卡上运行。安装 FreeVideo 时勾选 Prism 会自动下载,无需手动获取。

  • int8/(46.6 GiB):所有档位共用。W8A8 INT8 权重(128 维 Hadamard 旋转,从原版权重量化),轻量档使用的 LightX2V Wan2.2-I2V 260412 蒸馏(未合并的秩 256 LoRA),UMT5-XXL 文本编码器(bf16),Wan2.1 VAE 和 DAC 音频解码器。
  • bf16/(60.8 GiB):可选,仅极致档使用的原版 bf16 权重;文本编码器、分词器和 VAE 读取 int8/ 中的文件。

四个档位:轻量(蒸馏 8 步,H200 约 6 分钟)、标准(原版权重 20 步,15.5 分钟)、精细(30 步,23 分钟)、极致(官方 50 步,bf16 加精确注意力,62.5 分钟)。

许可证:Prism 权重与代码为 MIT(腾讯);MOVA、Wan2.1/Wan2.2(含 UMT5-XXL 文本编码器与 Wan VAE)和 LightX2V 蒸馏模型为 Apache-2.0;Descript Audio Codec 为 MIT。详见 LICENSE。

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for FreeVideoX/Prism-FreeVideo-Preview

Finetuned
(2)
this model