Spaces:
Running on Zero
Running on Zero
multimodalart HF Staff
ID-V2V ZeroGPU demo: SAM3 + foreground-on-gray VACE + Wan2.1 I2V-14B DiT
49bc52e verified A newer version of the Gradio SDK is available: 6.24.0
metadata
title: ID-V2V
emoji: πͺ
colorFrom: pink
colorTo: purple
sdk: gradio
sdk_version: 6.22.0
app_file: app.py
python_version: '3.12'
startup_duration_timeout: 1h
pinned: false
license: apache-2.0
short_description: Identity-preserving video restyle & relight
models:
- Eyeline-Labs/ID-V2V
- Wan-AI/Wan2.1-T2V-14B
- Wan-AI/Wan2.1-I2V-14B-480P
- facebook/sam3
tags:
- video-to-video
- video-editing
- identity-preservation
- wan
- vace
ID-V2V
Demo of Eyeline-Labs/ID-V2V β restyle or relight a video while keeping the people in it recognisable.
Inputs: a source video (supplies the motion), a stylized first frame (supplies the target look and identity), and a prompt.
Pipeline
Ported from the authors' reference implementation (github.com/Eyeline-Labs/ID-V2V):
- SAM3 promptable concept segmentation (
"person") tracks the subjects across the clip; masks are cleaned with the repo's Secret Panda routine (hole-fill β close β bridge β hole-fill), per object and then on the union. - The masked pixels are composited over gray 127 β this foreground-on-gray video is the single VACE control stream ID-V2V was trained on.
- Wan2.1 I2V-14B DiT + VACE ControlNet, both taken from the finetuned
idv2v.pth, run with the stylized frame as the I2V anchor and as the SVI anti-drift reference pad (ref_pad_num = -1),cfg_scale = 5,vace_scale = 1.
The authors' diffsynth fork (their custom multi-control VACE + SVI WanVideoPipeline) is
vendored under diffsynth/.
Deviations from the reference, and why
| Reference | Here | Reason |
|---|---|---|
| 1280Γ720 | 832Γ480 (the repo's own CLI default) | 720p is ~2.5Γ the tokens; it does not fit one ZeroGPU slot |
| 81 frames/clip, multi-clip chaining for long videos | 17/33/49 frames, single clip | each extra clip is a full extra denoise pass |
| Consecutive source frames | selectable source-frame stride (default 2, output fps divided to match) | a 33-frame clip then still covers the whole action |
| SAM3 at source resolution, then crop/resize | crop/resize to 832Γ480, then SAM3 | same masks, much faster segmentation |
Secret Panda morphology via scipy.ndimage |
same operations via OpenCV (border_value=0) |
identical results, ~10Γ faster |
| bf16 weights with CPU offload across 8ΓA100 | bf16 weights fully resident on one 96 GB ZeroGPU slot | no offload on ZeroGPU |
idv2v.pth ships as fp32 (β78 GB); it is memory-mapped, cast to bf16 (β39 GB), and the
fp32 blob is deleted before ZeroGPU packs the weights.
Credits
- Model & code: Eyeline Labs β paper
- Base model: Wan-AI/Wan2.1-I2V-14B-720P + VACE
- Segmentation: facebook/sam3
- Example clips are the authors' own
test_samples/.