---
title: InstructAV2AV
emoji: ๐ฌ
colorFrom: indigo
colorTo: purple
sdk: gradio
sdk_version: 6.8.0
python_version: "3.10"
app_file: app.py
pinned: false
license: apache-2.0
suggested_hardware: zero-a10g
suggested_storage: large
---
InstructAV2AV: Instruction-Guided Audio-Video Joint Editing

Haojie Zheng
1, 2, Yixin Yang
2 , Siqi Yang
2, Shuchen Weng
ย 1,2* , Boxin Shi
2*
* Corresponding authors
1 BAAI,
2 Peking University
---
## ๐ฅ Video Demo
|
|
|
For the best experience, please enable audio.
|
## ๐
Todo List
We are working hard to deliver the components as soon as possible. Here is our roadmap for the upcoming release:
- [x] **Inference Code** (Including pre-trained weights and pipeline scripts)
- [x] **InsAVE Dataset** (The complete dataset used for training and evaluation)
- [x] **Training Scripts** (Codes for reproducing or fine-tuning our model)
## โ๏ธ Installation
For CUDA 12.1, you can install the dependencies with the following commands.
```bash
git clone https://github.com/suimuc/InstructAV2AV.git
cd InstructAV2AV
conda create -n instructav2av python=3.10 -y
conda activate instructav2av
pip install torch==2.8.0 torchvision==0.23.0 torchaudio==2.8.0
pip install -r requirements.txt
pip install flash-attn==2.8.3.post1 --no-build-isolation
pip install -e .
```
## ๐ฅ Download Pretrained Models
Using `huggingface-cli` to download the models:
```bash
# InstructAV2AV Model
huggingface-cli download suimu/InstructAV2AV \
--local-dir ckpts/InstructAV2AV
# Text Encoder & Video VAE
huggingface-cli download Wan-AI/Wan2.2-TI2V-5B \
--include \
"google/*" \
"models_t5_umt5-xxl-enc-bf16.pth" \
"Wan2.2_VAE.pth" \
--local-dir ckpts/Wan2.2-TI2V-5B
# Audio VAE
huggingface-cli download hkchengrex/MMAudio \
--include \
"ext_weights/best_netG.pt" \
"ext_weights/v1-16.pth" \
--local-dir ckpts/MMAudio
```
## ๐ฎ Inference
### Command-Line
The unified inference entry point is `scripts/edit.py`. The source video must contain an audio track unless a separate file is provided with `--source-audio`.
```bash
python scripts/edit.py \
--source-video assets/input.mp4 \
--instruction "Keep the personโs identity and change the spoken words to This is more than just art, itโs a statement.." \
--finetune-path ckpts/InstructAV2AV/clone_id_voice.safetensors \
--output outputs/edited.mp4
```
To achieve optimal editing quality and better text-to-video alignment, you can enhance the instruction by appending a detailed description of the final edited video immediately after your editing command.
```bash
python scripts/edit.py \
--source-video assets/input1.mp4 \
--instruction "Change the man into a young woman with brown hair, wearing a gray blazer over a light pink top and a necklace with a heart-shaped pendant, and saying, I really think we should give it another chance.. A young woman with long, wavy brown hair and fair skin is engaged in a conversation. She is wearing a gray blazer over a light pink top and has a necklace with a heart-shaped pendant. Her facial expressions change throughout the sequence, showing a range of emotions that suggest she is either explaining something earnestly or reacting to a conversation, and says I really think we should give it another chance. The setting appears to be indoors, with a dimly lit, blurred background that suggests a social environment, possibly a bar or restaurant. The focus remains on the woman's face, capturing her reactions and engagement in the dialogue." \
--finetune-path ckpts/InstructAV2AV/general.safetensors \
--output outputs/edited.mp4
```
InstructAV2AV provides multiple task-specific checkpoints. All checkpoints use the same model architecture, but each checkpoint is fine-tuned for a different editing type. Select the checkpoint that best matches your instruction and provide its path through `--finetune-path`.
| Checkpoint | Editing Type | Description | Example Instruction |
| ---------------- | ------------------------ | ------------------------------------------------------------ | ------------------------------------------------------------ |
| `general` | General Edit | Flexible editing of appearance, scenes, actions, speech, and sound. | `Make the horse dark brown with a white saddle.` |
| `insertion` | Content Insertion | Add an object or other content to the source video. | `Add a dark vintage sedan driving from the right to the left.` |
| `removal` | Content Removal | Remove an object and its associated audiovisual content from the source video. | `Remove the chipmunk standing on the stone surface among the peanuts.` |
| `clone_id` | Identity Cloning | Preserve a person's visual identity while editing other visual or audio attributes. | `Keep the person's appearance, change the timbre to a man, and change the spoken words to I understand, but I think we need to consider..` |
| `clone_voice` | Voice Cloning | Preserve the speaker's timbre while editing the video or spoken content. | `Keep the timbre, change the person's appearance, and change the spoken words to I came here to tell you that you should go..` |
| `clone_id_voice` | Identity + Voice Cloning | Preserve both the person's visual identity and the speaker's timbre. | `Keep the person's identity and voice, and change the spoken words to This is more than just art, it's a statement..` |
### Gradio
After downloading all six InstructAV2AV checkpoints and the dependency weights, launch the demo with:
```bash
python scripts/demo.py --share
```
### Hugging Face Space
This repository is ready to upload directly to a native **Gradio Space**. Select **ZeroGPU** as the Space hardware and push the repository contents; Hugging Face launches `app.py` automatically.
At process startup, the Space downloads the shared Wan/MMAudio weights and the default `general` checkpoint. When the page opens, a `spaces.GPU` callback initializes the default engine on ZeroGPU before the user submits an edit. Other task-specific checkpoints remain lazy and are downloaded when selected. The uploaded source video must contain an audio track.
Optional Space variables and secrets:
| Name | Purpose |
| --- | --- |
| `HF_TOKEN` | Hub token, only needed if a dependency repository requires authentication. |
| `HF_HOME` | Override the Hub cache directory. When writable `/data` storage is attached, the app uses `/data/.huggingface` automatically. |
| `INSTRUCTAV2AV_MODEL_HOME` | Override the generated model-layout directory. It defaults to a subdirectory of `HF_HOME`. |
| `INSTRUCTAV2AV_CPU_OFFLOAD` | Set to `1` to enable CPU offload when GPU memory is constrained. Default: `0` for ZeroGPU. |
| `INSTRUCTAV2AV_EAGER_DOWNLOAD` | Set to `0` to skip startup downloading of the shared weights and `general` checkpoint. Default: `1`. |
| `INSTRUCTAV2AV_ZEROGPU_DURATION` | Requested ZeroGPU allocation duration in seconds. Default: `300`. |
| `GRADIO_MAX_FILE_SIZE` | Maximum upload size. Default: `500mb`. |
## ๐ฆ InsAVE-80K Dataset
The training dataset is available at [suimu/InsAVE-80K](https://huggingface.co/datasets/suimu/InsAVE-80K).
## ๐ Training
### Training manifest format
The trainer accepts CSV, JSON, or JSONL manifests with the following canonical fields:
| Field | Description |
| -------------- | :-----------------: |
| `source_video` | Original video |
| `source_audio` | Original audio |
| `target_video` | Edited video |
| `target_audio` | Edited audio |
| `instruction` | Editing instruction |
### Training scripts
The unified training entry point is `scripts/train.py`, with configuration in `ovi/configs/train/train_av_edit.yaml`.
Before training, update at least these entries:
```yaml
ckpt_dir: ./ckpts
finetune_path: ./ckpts/InstructAV2AV/general.safetensors
dataset:
metadata_path: ./data/InsAVE-80K/path/to/manifest.csv
```
The manifest and initialization checkpoint can also be overridden from the command line:
```bash
accelerate launch \
--config_file ovi/configs/train/accelerate_config.yaml \
scripts/train.py \
--config-file ovi/configs/train/train_av_edit.yaml \
--data-manifest data/InsAVE-80K/path/to/manifest.csv \
--finetune-path ckpts/InstructAV2AV/general.safetensors \
--output-dir outputs/train_av_edit
```
The provided Accelerate configuration uses DeepSpeed ZeRO-2 and is configured for eight processes. Change `num_processes` in `ovi/configs/train/accelerate_config.yaml` to match the number of available GPUs.
The training objective is selected automatically from the modality flags:
| `has_video` | `has_audio` | Training objective |
| ----------: | ----------: | ------------------------- |
| `true` | `true` | Joint audio-video editing |
| `true` | `false` | Video-only editing |
| `false` | `true` | Audio-only editing |
## ๐ Acknowledgements
This project builds on the following open-source projects:
- [Ovi](https://github.com/character-ai/Ovi) for the audio-video fusion backbone.
- [Wan2.2](https://github.com/Wan-Video/Wan2.2) for the video backbone components,
UMT5 text encoder, and video VAE.
- [MMAudio](https://github.com/hkchengrex/MMAudio) for the audio VAE and vocoder.
We thank the authors and contributors of these projects for releasing their work.
## โ๏ธ Citation
If you find our work or dataset helpful for your research, please consider citing our paper:
```bibtex
@article{instructav2av2026,
title={InstructAV2AV: Instruction-Guided Audio-Video Joint Editing},
author={Zheng, Haojie and Yang, Yixin and Yang, Siqi and Weng, Shuchen and Shi, Boxin},
journal={arXiv preprint arXiv:2605.18467},
year={2026}
}
```