--- title: InstructAV2AV emoji: ๐ŸŽฌ colorFrom: indigo colorTo: purple sdk: gradio sdk_version: 6.8.0 python_version: "3.10" app_file: app.py pinned: false license: apache-2.0 suggested_hardware: zero-a10g suggested_storage: large ---

InstructAV2AV: Instruction-Guided Audio-Video Joint Editing

Haojie Zheng1, 2, Yixin Yang2 , Siqi Yang2, Shuchen Wengย 1,2* , Boxin Shi 2*
* Corresponding authors
1 BAAI, 2 Peking University
--- ## ๐ŸŽฅ Video Demo
For the best experience, please enable audio.
## ๐Ÿ“… Todo List We are working hard to deliver the components as soon as possible. Here is our roadmap for the upcoming release: - [x] **Inference Code** (Including pre-trained weights and pipeline scripts) - [x] **InsAVE Dataset** (The complete dataset used for training and evaluation) - [x] **Training Scripts** (Codes for reproducing or fine-tuning our model) ## โš™๏ธ Installation For CUDA 12.1, you can install the dependencies with the following commands. ```bash git clone https://github.com/suimuc/InstructAV2AV.git cd InstructAV2AV conda create -n instructav2av python=3.10 -y conda activate instructav2av pip install torch==2.8.0 torchvision==0.23.0 torchaudio==2.8.0 pip install -r requirements.txt pip install flash-attn==2.8.3.post1 --no-build-isolation pip install -e . ``` ## ๐Ÿ“ฅ Download Pretrained Models Using `huggingface-cli` to download the models: ```bash # InstructAV2AV Model huggingface-cli download suimu/InstructAV2AV \ --local-dir ckpts/InstructAV2AV # Text Encoder & Video VAE huggingface-cli download Wan-AI/Wan2.2-TI2V-5B \ --include \ "google/*" \ "models_t5_umt5-xxl-enc-bf16.pth" \ "Wan2.2_VAE.pth" \ --local-dir ckpts/Wan2.2-TI2V-5B # Audio VAE huggingface-cli download hkchengrex/MMAudio \ --include \ "ext_weights/best_netG.pt" \ "ext_weights/v1-16.pth" \ --local-dir ckpts/MMAudio ``` ## ๐ŸŽฎ Inference ### Command-Line The unified inference entry point is `scripts/edit.py`. The source video must contain an audio track unless a separate file is provided with `--source-audio`. ```bash python scripts/edit.py \ --source-video assets/input.mp4 \ --instruction "Keep the personโ€™s identity and change the spoken words to This is more than just art, itโ€™s a statement.." \ --finetune-path ckpts/InstructAV2AV/clone_id_voice.safetensors \ --output outputs/edited.mp4 ``` To achieve optimal editing quality and better text-to-video alignment, you can enhance the instruction by appending a detailed description of the final edited video immediately after your editing command. ```bash python scripts/edit.py \ --source-video assets/input1.mp4 \ --instruction "Change the man into a young woman with brown hair, wearing a gray blazer over a light pink top and a necklace with a heart-shaped pendant, and saying, I really think we should give it another chance.. A young woman with long, wavy brown hair and fair skin is engaged in a conversation. She is wearing a gray blazer over a light pink top and has a necklace with a heart-shaped pendant. Her facial expressions change throughout the sequence, showing a range of emotions that suggest she is either explaining something earnestly or reacting to a conversation, and says I really think we should give it another chance. The setting appears to be indoors, with a dimly lit, blurred background that suggests a social environment, possibly a bar or restaurant. The focus remains on the woman's face, capturing her reactions and engagement in the dialogue." \ --finetune-path ckpts/InstructAV2AV/general.safetensors \ --output outputs/edited.mp4 ``` InstructAV2AV provides multiple task-specific checkpoints. All checkpoints use the same model architecture, but each checkpoint is fine-tuned for a different editing type. Select the checkpoint that best matches your instruction and provide its path through `--finetune-path`. | Checkpoint | Editing Type | Description | Example Instruction | | ---------------- | ------------------------ | ------------------------------------------------------------ | ------------------------------------------------------------ | | `general` | General Edit | Flexible editing of appearance, scenes, actions, speech, and sound. | `Make the horse dark brown with a white saddle.` | | `insertion` | Content Insertion | Add an object or other content to the source video. | `Add a dark vintage sedan driving from the right to the left.` | | `removal` | Content Removal | Remove an object and its associated audiovisual content from the source video. | `Remove the chipmunk standing on the stone surface among the peanuts.` | | `clone_id` | Identity Cloning | Preserve a person's visual identity while editing other visual or audio attributes. | `Keep the person's appearance, change the timbre to a man, and change the spoken words to I understand, but I think we need to consider..` | | `clone_voice` | Voice Cloning | Preserve the speaker's timbre while editing the video or spoken content. | `Keep the timbre, change the person's appearance, and change the spoken words to I came here to tell you that you should go..` | | `clone_id_voice` | Identity + Voice Cloning | Preserve both the person's visual identity and the speaker's timbre. | `Keep the person's identity and voice, and change the spoken words to This is more than just art, it's a statement..` | ### Gradio After downloading all six InstructAV2AV checkpoints and the dependency weights, launch the demo with: ```bash python scripts/demo.py --share ``` ### Hugging Face Space This repository is ready to upload directly to a native **Gradio Space**. Select **ZeroGPU** as the Space hardware and push the repository contents; Hugging Face launches `app.py` automatically. At process startup, the Space downloads the shared Wan/MMAudio weights and the default `general` checkpoint. When the page opens, a `spaces.GPU` callback initializes the default engine on ZeroGPU before the user submits an edit. Other task-specific checkpoints remain lazy and are downloaded when selected. The uploaded source video must contain an audio track. Optional Space variables and secrets: | Name | Purpose | | --- | --- | | `HF_TOKEN` | Hub token, only needed if a dependency repository requires authentication. | | `HF_HOME` | Override the Hub cache directory. When writable `/data` storage is attached, the app uses `/data/.huggingface` automatically. | | `INSTRUCTAV2AV_MODEL_HOME` | Override the generated model-layout directory. It defaults to a subdirectory of `HF_HOME`. | | `INSTRUCTAV2AV_CPU_OFFLOAD` | Set to `1` to enable CPU offload when GPU memory is constrained. Default: `0` for ZeroGPU. | | `INSTRUCTAV2AV_EAGER_DOWNLOAD` | Set to `0` to skip startup downloading of the shared weights and `general` checkpoint. Default: `1`. | | `INSTRUCTAV2AV_ZEROGPU_DURATION` | Requested ZeroGPU allocation duration in seconds. Default: `300`. | | `GRADIO_MAX_FILE_SIZE` | Maximum upload size. Default: `500mb`. | ## ๐Ÿ“ฆ InsAVE-80K Dataset The training dataset is available at [suimu/InsAVE-80K](https://huggingface.co/datasets/suimu/InsAVE-80K). ## ๐Ÿš€ Training ### Training manifest format The trainer accepts CSV, JSON, or JSONL manifests with the following canonical fields: | Field | Description | | -------------- | :-----------------: | | `source_video` | Original video | | `source_audio` | Original audio | | `target_video` | Edited video | | `target_audio` | Edited audio | | `instruction` | Editing instruction | ### Training scripts The unified training entry point is `scripts/train.py`, with configuration in `ovi/configs/train/train_av_edit.yaml`. Before training, update at least these entries: ```yaml ckpt_dir: ./ckpts finetune_path: ./ckpts/InstructAV2AV/general.safetensors dataset: metadata_path: ./data/InsAVE-80K/path/to/manifest.csv ``` The manifest and initialization checkpoint can also be overridden from the command line: ```bash accelerate launch \ --config_file ovi/configs/train/accelerate_config.yaml \ scripts/train.py \ --config-file ovi/configs/train/train_av_edit.yaml \ --data-manifest data/InsAVE-80K/path/to/manifest.csv \ --finetune-path ckpts/InstructAV2AV/general.safetensors \ --output-dir outputs/train_av_edit ``` The provided Accelerate configuration uses DeepSpeed ZeRO-2 and is configured for eight processes. Change `num_processes` in `ovi/configs/train/accelerate_config.yaml` to match the number of available GPUs. The training objective is selected automatically from the modality flags: | `has_video` | `has_audio` | Training objective | | ----------: | ----------: | ------------------------- | | `true` | `true` | Joint audio-video editing | | `true` | `false` | Video-only editing | | `false` | `true` | Audio-only editing | ## ๐Ÿ™ Acknowledgements This project builds on the following open-source projects: - [Ovi](https://github.com/character-ai/Ovi) for the audio-video fusion backbone. - [Wan2.2](https://github.com/Wan-Video/Wan2.2) for the video backbone components, UMT5 text encoder, and video VAE. - [MMAudio](https://github.com/hkchengrex/MMAudio) for the audio VAE and vocoder. We thank the authors and contributors of these projects for releasing their work. ## โœ’๏ธ Citation If you find our work or dataset helpful for your research, please consider citing our paper: ```bibtex @article{instructav2av2026, title={InstructAV2AV: Instruction-Guided Audio-Video Joint Editing}, author={Zheng, Haojie and Yang, Yixin and Yang, Siqi and Weng, Shuchen and Shi, Boxin}, journal={arXiv preprint arXiv:2605.18467}, year={2026} } ```