Spaces:
Running on Zero
A newer version of the Gradio SDK is available: 6.22.0
title: InstructAV2AV
emoji: ๐ฌ
colorFrom: indigo
colorTo: purple
sdk: gradio
sdk_version: 6.8.0
python_version: '3.10'
app_file: app.py
pinned: false
license: apache-2.0
suggested_hardware: zero-a10g
suggested_storage: large
InstructAV2AV: Instruction-Guided Audio-Video Joint Editing
Haojie Zheng1, 2, Yixin Yang2 , Siqi Yang2, Shuchen Weng 1,2* , Boxin Shi 2*
* Corresponding authors
1 BAAI, 2 Peking University
๐ฅ Video Demo
| For the best experience, please enable audio. |
๐ Todo List
We are working hard to deliver the components as soon as possible. Here is our roadmap for the upcoming release:
- Inference Code (Including pre-trained weights and pipeline scripts)
- InsAVE Dataset (The complete dataset used for training and evaluation)
- Training Scripts (Codes for reproducing or fine-tuning our model)
โ๏ธ Installation
For CUDA 12.1, you can install the dependencies with the following commands.
git clone https://github.com/suimuc/InstructAV2AV.git
cd InstructAV2AV
conda create -n instructav2av python=3.10 -y
conda activate instructav2av
pip install torch==2.8.0 torchvision==0.23.0 torchaudio==2.8.0
pip install -r requirements.txt
pip install flash-attn==2.8.3.post1 --no-build-isolation
pip install -e .
๐ฅ Download Pretrained Models
Using huggingface-cli to download the models:
# InstructAV2AV Model
huggingface-cli download suimu/InstructAV2AV \
--local-dir ckpts/InstructAV2AV
# Text Encoder & Video VAE
huggingface-cli download Wan-AI/Wan2.2-TI2V-5B \
--include \
"google/*" \
"models_t5_umt5-xxl-enc-bf16.pth" \
"Wan2.2_VAE.pth" \
--local-dir ckpts/Wan2.2-TI2V-5B
# Audio VAE
huggingface-cli download hkchengrex/MMAudio \
--include \
"ext_weights/best_netG.pt" \
"ext_weights/v1-16.pth" \
--local-dir ckpts/MMAudio
๐ฎ Inference
Command-Line
The unified inference entry point is scripts/edit.py. The source video must contain an audio track unless a separate file is provided with --source-audio.
python scripts/edit.py \
--source-video assets/input.mp4 \
--instruction "Keep the personโs identity and change the spoken words to <S>This is more than just art, itโs a statement.<E>." \
--finetune-path ckpts/InstructAV2AV/clone_id_voice.safetensors \
--output outputs/edited.mp4
To achieve optimal editing quality and better text-to-video alignment, you can enhance the instruction by appending a detailed description of the final edited video immediately after your editing command.
python scripts/edit.py \
--source-video assets/input1.mp4 \
--instruction "Change the man into a young woman with brown hair, wearing a gray blazer over a light pink top and a necklace with a heart-shaped pendant, and saying, <S>I really think we should give it another chance.<E>. A young woman with long, wavy brown hair and fair skin is engaged in a conversation. She is wearing a gray blazer over a light pink top and has a necklace with a heart-shaped pendant. Her facial expressions change throughout the sequence, showing a range of emotions that suggest she is either explaining something earnestly or reacting to a conversation, and says <S>I really think we should give it another chance.<E> The setting appears to be indoors, with a dimly lit, blurred background that suggests a social environment, possibly a bar or restaurant. The focus remains on the woman's face, capturing her reactions and engagement in the dialogue." \
--finetune-path ckpts/InstructAV2AV/general.safetensors \
--output outputs/edited.mp4
InstructAV2AV provides multiple task-specific checkpoints. All checkpoints use the same model architecture, but each checkpoint is fine-tuned for a different editing type. Select the checkpoint that best matches your instruction and provide its path through --finetune-path.
| Checkpoint | Editing Type | Description | Example Instruction |
|---|---|---|---|
general |
General Edit | Flexible editing of appearance, scenes, actions, speech, and sound. | Make the horse dark brown with a white saddle. |
insertion |
Content Insertion | Add an object or other content to the source video. | Add a dark vintage sedan driving from the right to the left. |
removal |
Content Removal | Remove an object and its associated audiovisual content from the source video. | Remove the chipmunk standing on the stone surface among the peanuts. |
clone_id |
Identity Cloning | Preserve a person's visual identity while editing other visual or audio attributes. | Keep the person's appearance, change the timbre to a man, and change the spoken words to <S>I understand, but I think we need to consider.<E>. |
clone_voice |
Voice Cloning | Preserve the speaker's timbre while editing the video or spoken content. | Keep the timbre, change the person's appearance, and change the spoken words to <S>I came here to tell you that you should go.<E>. |
clone_id_voice |
Identity + Voice Cloning | Preserve both the person's visual identity and the speaker's timbre. | Keep the person's identity and voice, and change the spoken words to <S>This is more than just art, it's a statement.<E>. |
Gradio
After downloading all six InstructAV2AV checkpoints and the dependency weights, launch the demo with:
python scripts/demo.py --share
Hugging Face Space
This repository is ready to upload directly to a native Gradio Space. Select ZeroGPU as the Space hardware and push the repository contents; Hugging Face launches app.py automatically.
At process startup, the Space downloads the shared Wan/MMAudio weights and the default general checkpoint. When the page opens, a spaces.GPU callback initializes the default engine on ZeroGPU before the user submits an edit. Other task-specific checkpoints remain lazy and are downloaded when selected. The uploaded source video must contain an audio track.
Optional Space variables and secrets:
| Name | Purpose |
|---|---|
HF_TOKEN |
Hub token, only needed if a dependency repository requires authentication. |
HF_HOME |
Override the Hub cache directory. When writable /data storage is attached, the app uses /data/.huggingface automatically. |
INSTRUCTAV2AV_MODEL_HOME |
Override the generated model-layout directory. It defaults to a subdirectory of HF_HOME. |
INSTRUCTAV2AV_CPU_OFFLOAD |
Set to 1 to enable CPU offload when GPU memory is constrained. Default: 0 for ZeroGPU. |
INSTRUCTAV2AV_EAGER_DOWNLOAD |
Set to 0 to skip startup downloading of the shared weights and general checkpoint. Default: 1. |
INSTRUCTAV2AV_ZEROGPU_DURATION |
Requested ZeroGPU allocation duration in seconds. Default: 300. |
GRADIO_MAX_FILE_SIZE |
Maximum upload size. Default: 500mb. |
๐ฆ InsAVE-80K Dataset
The training dataset is available at suimu/InsAVE-80K.
๐ Training
Training manifest format
The trainer accepts CSV, JSON, or JSONL manifests with the following canonical fields:
| Field | Description |
|---|---|
source_video |
Original video |
source_audio |
Original audio |
target_video |
Edited video |
target_audio |
Edited audio |
instruction |
Editing instruction |
Training scripts
The unified training entry point is scripts/train.py, with configuration in ovi/configs/train/train_av_edit.yaml.
Before training, update at least these entries:
ckpt_dir: ./ckpts
finetune_path: ./ckpts/InstructAV2AV/general.safetensors
dataset:
metadata_path: ./data/InsAVE-80K/path/to/manifest.csv
The manifest and initialization checkpoint can also be overridden from the command line:
accelerate launch \
--config_file ovi/configs/train/accelerate_config.yaml \
scripts/train.py \
--config-file ovi/configs/train/train_av_edit.yaml \
--data-manifest data/InsAVE-80K/path/to/manifest.csv \
--finetune-path ckpts/InstructAV2AV/general.safetensors \
--output-dir outputs/train_av_edit
The provided Accelerate configuration uses DeepSpeed ZeRO-2 and is configured for eight processes. Change num_processes in ovi/configs/train/accelerate_config.yaml to match the number of available GPUs.
The training objective is selected automatically from the modality flags:
has_video |
has_audio |
Training objective |
|---|---|---|
true |
true |
Joint audio-video editing |
true |
false |
Video-only editing |
false |
true |
Audio-only editing |
๐ Acknowledgements
This project builds on the following open-source projects:
- Ovi for the audio-video fusion backbone.
- Wan2.2 for the video backbone components, UMT5 text encoder, and video VAE.
- MMAudio for the audio VAE and vocoder.
We thank the authors and contributors of these projects for releasing their work.
โ๏ธ Citation
If you find our work or dataset helpful for your research, please consider citing our paper:
@article{instructav2av2026,
title={InstructAV2AV: Instruction-Guided Audio-Video Joint Editing},
author={Zheng, Haojie and Yang, Yixin and Yang, Siqi and Weng, Shuchen and Shi, Boxin},
journal={arXiv preprint arXiv:2605.18467},
year={2026}
}