--- license: mit library_name: harp-codec pipeline_tag: audio-to-audio tags: - audio - neural-audio-codec - audio-compression - residual-vector-quantization - rvq - interspeech-2026 --- # HARP: Harmonic-Aware Residual Partitioning for Neural Audio Codecs HARP is a neural audio codec that partitions residual vector quantization (RVQ) across harmonically meaningful frequency bands, giving high-quality, variable-bitrate audio compression from a single model. A harmonic-aware partitioning distributes codebook capacity across perceptually meaningful bands, so one model serves multiple bitrates by decoding a growing number of codebook groups. - 📄 **Paper:** [arXiv:2607.16657](https://arxiv.org/abs/2607.16657) (Interspeech 2026, Oral) - 💻 **Code:** https://github.com/QiaoyuYang/harp-codec - 🔊 **Audio examples:** https://qiaoyuyang.github.io/harp-codec/ ## Model details - **Architecture:** convolutional encoder/decoder with grouped RVQ. - **Codebooks:** 9 codebooks × 1024 entries (10 bits each), codebook dim 8. - **Band groups:** 4 groups with a 3-2-2-2 codebook split, ordered by perceptual importance: | Group | Band | Codebooks | |-------|------|-----------| | 0 | Bass (~0–1 kHz) | 3 | | 1 | Low-mid (~1–4 kHz) | 2 | | 2 | High-mid (~4–10 kHz) | 2 | | 3 | Treble (~10–22 kHz) | 2 | The bass group is always decoded; adding groups raises quality and bitrate. ## Bitrate tiers With 9 codebooks @ 1024 entries and a ~86 Hz frame rate: | Groups | Codebooks | Approx. bitrate | |--------|-----------|-----------------| | 1 | 3 | ~2.6 kbps | | 2 | 5 | ~4.3 kbps | | 3 | 7 | ~6.0 kbps | | 4 | 9 | ~7.7 kbps (full) | ## Files - `harp.ckpt` — slim, inference-ready weights (~293 MB, weights only). ## Usage Set up the [GitHub repo](https://github.com/QiaoyuYang/harp-codec) (`uv sync`), download this checkpoint, and reconstruct audio: ```bash uv pip install "huggingface_hub[cli]" hf download KelvinYang/harp-codec harp.ckpt --local-dir checkpoints python entry.py -i --input audio.wav --output recon.wav # full rate python entry.py -i --input audio.wav --n-groups 2 --output out.wav # lower bitrate (1..4) ``` Each run reports SI-SDR, multi-scale mel loss, LSD, and SNR. ## Citation ```bibtex @inproceedings{harp2026, title = {HARP: Harmonic-Aware Residual Partitioning for Neural Audio Codecs}, author = {Yang, Qiaoyu and He, Lixing and Deng, Binyue and Zhao, Weifeng}, booktitle = {Interspeech}, year = {2026} } ``` ## License MIT