--- language: - en - zh license: apache-2.0 tags: - music - vae - codec pipeline_tag: audio-to-audio --- # εar-VAE2 **Fourier is Frontier: Frequency-Aware Autoencoding for High-Fidelity Music Reconstruction** Kangdi Wang1 · Yusheng Dai2 · Jin Xu1† 1 Qwen Team, Alibaba    2 Monash University    Corresponding author [[Demo Page](https://eps-acoustic-revolution-lab.github.io/EAR_VAE2/)] - [[Paper](https://arxiv.org/abs/2608.19843)] - [[Codebase](https://github.com/Eps-Acoustic-Revolution-Lab/EAR_VAE2)] ---

εar-VAE2 Architecture

## Overview A spectral-domain music autoencoder compressing **48 kHz stereo** audio into a **128-dimensional** continuous latent sequence at **25 Hz** — a **1920× temporal downsampling** — through two frequency-aware components: **Spec-SnakeBeta** (per-bin periodic activation with log-frequency initialization) and a **Duplex-Aware Refiner** (band-specific magnitude/phase correction motivated by psychoacoustic masking). ## ✨ Highlights - 🎵 **Complex spectral domain** — operates on STFT real/imag channels, not raw waveform - 🧬 **Spec-SnakeBeta** — per-(channel, frequency-bin) periodic activation with log-frequency initialization; each bin learns its own oscillatory bias - 🎛️ **Duplex-Aware Refiner** — band-specific mag/phase correction following psychoacoustic dominance (phase-only < 1.5 kHz, joint mid-band, mag-only > 4 kHz) - 📊 **1920× compression** — 48 kHz stereo → 128-d × 25 Hz continuous latent - 🏆 **SOTA reconstruction** on Song Describer Dataset across spectral metrics ## Main Results Reconstruction quality on **Song Describer Dataset** (546 full tracks, 48 kHz stereo): | System | SI-SDR ↑ | STFT Dist ↓ | Mel Dist ↓ | CCPC ↑ | |--------|:---:|:---:|:---:|:---:| | εar-VAE | 12.4 | 0.880 | 0.509 | 0.973 | | SA-Open | 6.7 | 1.016 | 0.612 | 0.933 | | Levo 2 | 8.1 | 0.971 | 0.599 | 0.947 | | SAME-L | 12.5 | 0.986 | 0.539 | 0.970 | | **εar-VAE2 (base)** | 10.9 | 0.916 | 0.572 | 0.966 | | **εar-VAE2 (full)** | **11.3** | **0.870** | **0.461** | **0.973** | > εar-VAE2 (full) achieves the best spectral fidelity (STFT Dist, Mel Dist) among all systems while matching the phase coherence (CCPC) of the εar-VAE baseline. ## Spec-SnakeBeta

Spec-SnakeBeta activation visualization

*Per-(channel, frequency-bin) periodic activation with log-scale parameterization. Low-frequency bins stay near-identity; high-frequency bins become progressively oscillatory — providing a physically motivated inductive bias for spectral processing.* ## Input Representation

Five-paradigm input representation comparison

*Complex STFT preserves organized high-frequency harmonic structure (panel A) where the same-backbone waveform-patch paradigm degrades (panel B). The spectral domain provides a physical frequency-axis inductive bias unavailable to waveform methods.* ## Installation ```bash # Clone the repository git clone https://github.com/Eps-Acoustic-Revolution-Lab/EAR_VAE2.git cd EAR_VAE2 # Install dependencies pip install -r requirements.txt # Download config + pretrained weights from the Hub huggingface-cli download earlab/EAR_VAE2 --local-dir checkpoints/ ``` ## Usage ### Python API ```python import json import torch from huggingface_hub import hf_hub_download from ear_vae2 import EarVAE2 REPO_ID = "earlab/EAR_VAE2" # Resolve config and weights from the Hub (cached locally after the first call) config_path = hf_hub_download(REPO_ID, "config.json") ckpt_path = hf_hub_download(REPO_ID, "weights/ear_vae2.pt") with open(config_path) as f: config = json.load(f)["model"]["gen"]["config"] model = EarVAE2(config) ckpt = torch.load(ckpt_path, map_location="cpu") model.load_state_dict(ckpt["gen"] if "gen" in ckpt else ckpt) model.eval().cuda() # Encode & decode audio = torch.randn(1, 2, 48000 * 10).cuda() # 10s stereo @ 48kHz audio_padded, orig_len = model.preprocess_audio(audio) latents = model.encode_audio(audio_padded, chunked=True, chunk_size=512, overlap=16, deterministic=True) reconstructed = model.decode_audio(latents, chunked=True, chunk_size=512, overlap=16) reconstructed = reconstructed[:, :, :orig_len] ``` ### Command Line Using the `checkpoints/` directory populated in the Installation step: ```bash python inference.py \ --checkpoint checkpoints/weights/ear_vae2.pt \ --config checkpoints/config.json \ --input input.wav --output output.wav ``` ## Model Details | Config | Params (M) | Latent dim | Rate (Hz) | Compression | |--------|:---:|:---:|:---:|:---:| | Small (C0=64) | ~42.6 | 128 | 25 | 1920× | - **Sample rate**: 48 kHz stereo - **STFT**: 3840-point FFT, 1920-sample hop → 25 Hz frame rate - **Latent**: 128-d continuous (VAE with KL regularization) - **Refiner**: 12-layer banded Transformer (256-d, 1024 intermediate) > **⚠️ Note on open-source weights:** Due to data licensing constraints, the open-source model weights are **retrained on publicly available datasets** (not the full internal training corpus). Performance may differ from the numbers reported in the paper, which were obtained with the full-scale proprietary training data. --- ## Citation ```bibtex @misc{earvae2, title = {Fourier is Frontier: Frequency-Aware Autoencoding for High-Fidelity Music Reconstruction}, author = {Kangdi Wang and Yusheng Dai and Jin Xu}, year = {2026}, eprint = {2608.19843}, archivePrefix = {arXiv}, primaryClass = {cs.SD}, url = {https://arxiv.org/abs/2608.19843} } ``` --- ## Acknowledgements We gratefully acknowledge the following projects that inspired components of εar-VAE2: - [**BigVGAN**](https://github.com/NVIDIA/BigVGAN) — SnakeBeta periodic activation design - [**Vocos**](https://github.com/gemelo-ai/vocos) — ConvNeXt block architecture for spectral modeling - [**Stable Audio Tools**](https://github.com/Stability-AI/stable-audio-tools) — Training infrastructure and audio pipeline patterns --- ## License This project is licensed under the [Apache License 2.0](LICENSE).