File size: 2,205 Bytes
a7477f5
d661d4a
fed6c68
b8e6339
 
a7477f5
b8e6339
a7477f5
08d6eb2
 
d661d4a
 
 
 
 
a7477f5
d661d4a
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
---
title: Bernini Diffusers v2
emoji: 🗿
colorFrom: yellow
colorTo: purple
sdk: gradio
sdk_version: 6.15.0
app_file: app.py
pinned: false
license: apache-2.0
short_description: Reference-to-video with ByteDance Bernini-Diffusers-v2
python_version: "3.12"
startup_duration_timeout: 2h
models:
  - ByteDance/Bernini-Diffusers-v2
---

# Bernini-Diffusers-v2 — reference-to-video

Give it a handful of **reference images** (a subject, an outfit, a prop, a scene…) and a prompt
that points at them as `image0`, `image1`, … Bernini's Qwen2.5-VL planner reads the references
together with the instruction and *plans* a target visual embedding with a flow-matching head; the
Wan2.2-A14B MoE renderer (two 14 B DiTs, high-noise + low-noise) turns that plan into a video.

- Model: [`ByteDance/Bernini-Diffusers-v2`](https://huggingface.co/ByteDance/Bernini-Diffusers-v2)
- Code: [`bytedance/Bernini`](https://github.com/bytedance/Bernini)

## What this Space runs

The `r2v` (reference-to-video) task, matching the authors' `scripts/bernini_v2/run_r2v.sh`
one-for-one: `guidance_mode=vae_txt_vit_wapg`, `omega_txt=4.5`, `omega_tgt=1.5`, `omega_img=3.0`,
`omega_vid=1.0`, `omega_scale=0.75`, `planning_step=50`, `vit_denoising_step=1`, `vit_txt_cfg=1.2`,
`vit_img_cfg=1.0`, `flow_shift=5.0`, `max_image_size=842`, 16 fps, and the same system / negative
prompt.

The only deviation is the default clip length and step count (33 frames / 16 steps instead of
81 / 40), so a generation fits inside a single ZeroGPU slot — both are sliders under
**Advanced settings**. At the defaults a video takes about 4 minutes.

The released checkpoint is fp32 (~180 GB); it is loaded in bf16, which is the dtype the reference
pipeline computes in anyway.

## Credits

The `bernini/` package and the `veomni/` subset shipped alongside `app.py` are vendored from
[`bytedance/Bernini`](https://github.com/bytedance/Bernini) and
[`ByteDance-Seed/VeOmni`](https://github.com/ByteDance-Seed/VeOmni) (v0.1.11), both Apache-2.0,
because both declare `requires-python` ranges that exclude this runtime.

The example reference images in `examples/` are the authors' own r2v test case assets from
`bytedance/Bernini` (Apache-2.0).