--- library_name: pytorch datasets: - lemonweed6312/VGGSound-Duet-Mask tags: - audio-visual-localization - sound-source-localization - self-supervised-learning - multimodal - pytorch --- # SCAV Checkpoints Official pretrained checkpoints for **SCAV** from *Whence the Voice? Self-supervised Dual-source Audio-Visual Localisation via Selective Convergence* (ECCV 2026). [Project page](https://happy-new-bears.github.io/scav-project-page/) · [Paper](https://happy-new-bears.github.io/scav-project-page/Han_ECCV_2026_CAMERA_READY.pdf) · [Code](https://github.com/happy-new-bears/SCAV) · [VGGSound-Duet^Mask](https://huggingface.co/datasets/lemonweed6312/VGGSound-Duet-Mask) ## Model description Localising multiple simultaneous sound sources is challenging because source separation benefits from knowing source locations, while localisation benefits from separated audio. SCAV uses the **selective convergence** behaviour of contrastive audio-visual models in a progressive two-stage framework: 1. **Stage 1** learns dominant-source localisation from mixed audio using contrastive learning. 2. **Stage 2** uses the Stage 1 spatial prior to decouple visual and audio features and reveal the subdominant source. The method is self-supervised and does not require manual localisation annotations during training. The paper reports single-forward-pass localisation of both sources at 43.2 FPS on one NVIDIA A100 GPU. ## Files | File | Architecture | Purpose | |---|---|---| | `checkpoints/best-stage1.pth` | `FlowAttnHardWayModel` | Dominant-source localisation and generation of Stage 1 spatial-prior heatmaps | | `checkpoints/best-stage2.pth` | `ASLNet` | Progressive dual-source localisation using the Stage 1 spatial prior | The checkpoints are provided together because Stage 2 is part of the same progressive SCAV pipeline and depends on heatmaps produced by Stage 1. ## Download The repository is public and does not require an access token. ### Hugging Face CLI ```bash pip install -U "huggingface_hub[cli]" hf download lemonweed6312/SCAV --local-dir ./SCAV-weights ``` To download one checkpoint only: ```bash hf download lemonweed6312/SCAV checkpoints/best-stage1.pth \ --local-dir ./SCAV-weights ``` ### Python ```python from huggingface_hub import hf_hub_download stage1_path = hf_hub_download( repo_id="lemonweed6312/SCAV", filename="checkpoints/best-stage1.pth", ) stage2_path = hf_hub_download( repo_id="lemonweed6312/SCAV", filename="checkpoints/best-stage2.pth", ) ``` ## Requirements Use these checkpoints with the [official SCAV repository](https://github.com/happy-new-bears/SCAV): ```bash git clone https://github.com/happy-new-bears/SCAV.git cd SCAV ``` The main requirements are Python 3.8 or newer, PyTorch 1.12 or newer, NumPy, h5py, OpenCV, scikit-learn, matplotlib, tqdm, and FFmpeg. CUDA is recommended. See the code repository for data preprocessing and the complete runtime setup. ## Loading the checkpoints The official evaluation scripts instantiate the corresponding model, load the checkpoint with `torch.load`, and accept either a wrapped `model_state_dict` or a plain state dictionary: ```python import torch checkpoint = torch.load("/path/to/best-stage1.pth", map_location="cpu") state_dict = checkpoint.get("model_state_dict", checkpoint) # Instantiate FlowAttnHardWayModel exactly as in stage1_vggss/eval.py # or stage1_vggss/eval_seg.py, then: model.load_state_dict(state_dict) model.eval() ``` Use `FlowAttnHardWayModel` for `best-stage1.pth` and `ASLNet` for `best-stage2.pth`. Model definitions and constructor settings are maintained in the official code repository. ## Citation If you use these checkpoints, please cite: ```bibtex @inproceedings{hu2026scav, title = {Whence the Voice? Self-supervised Dual-source Audio-Visual Localisation via Selective Convergence}, author = {Hu, Han and Lin, Dongheng and Hou, Yuqi and Li, Haotian and Chang, Hyung Jin and Jiao, Jianbo}, booktitle = {European Conference on Computer Vision (ECCV)}, year = {2026} } ```