--- title: Vocal Remover emoji: 🎤 colorFrom: purple colorTo: pink sdk: gradio sdk_version: 6.19.0 app_file: app.py pinned: false license: mit --- # 🎤 Vocal Remover — Kim_Vocal_2 (MDX-Net ONNX) Separate **vocals** and **instrumental** from any song using the `Kim_Vocal_2.onnx` model — the best quality-per-byte ONNX vocal separator available. ## What it does - Upload any audio file (MP3, WAV, FLAC, M4A, OGG…) - Get back a clean **Vocals Only** stem and an **Instrumental** stem - See waveform + mel spectrogram visualizations for all three tracks - View proxy quality metrics (leakage dB, energy ratio) ## Model | | | |---|---| | **Model** | `Kim_Vocal_2.onnx` (MDX-Net / TFC-TDF U-Net) | | **Size** | ~67 MB | | **Vocal SDR** | ~8.9 dB on MVSep benchmark | | **Runtime** | ONNX Runtime (CPU) | | **Author** | Kimberley Jensen | ## Performance Running on free-tier CPU hardware. A typical 3-minute song takes **2–4 minutes** to process. The model is downloaded automatically on first use (~67 MB) and cached for the rest of the session. ## Settings | Setting | Default | Notes | |---|---|---| | Segment Size | 256 | Lower = less RAM; 512 = better quality | | Overlap | 0.25 | 0.5 = smoother transitions, 2× slower | | Enable Denoise | ✅ | Reduces post-separation artifacts | ## Updating the Gradio version If you fork this Space and get a build error, update `sdk_version` in this file to match the latest available Gradio version on HuggingFace Spaces: