--- license: mit library_name: candle tags: - voice-conversion - vocoder - real-time - low-latency - rust - candle - asmr language: - ja pipeline_tag: audio-to-audio --- # LightVC — model suite Own-weights model suite for real-time Japanese voice conversion (ASMR / バ美声). Inference is **pure Rust / [Candle](https://github.com/huggingface/candle)** (no Python runtime), target **E2E < 50 ms, CPU, causal**. Code: . License **MIT**. All weights are **100% own** (no distillation from other voice-conversion models; targets are real audio only). ## Components | component | folder | status | notes | |---|---|---|---| | **vocoder (freeC)** | `vocoder/` | ✅ released | low-latency neural vocoder, ear-parity with BigVGAN | | content encoder (G-enc) | `content-encoder/` | ⬜ planned | causal HuBERT-distilled, leakage-suppressed | | VC generator (AdaIN + CIPT) | `vc/` | ⬜ planned | mel generator, multi-reference moe-factor compositor | | prosody / s_art | `prosody/` | ⬜ planned | articulation clone, prosody policy | | voices | `voices/` | ⬜ planned | target / reference embeddings for the compositor | Download the whole suite or one component: ```bash huggingface-cli download mus8tte/lightvc --local-dir models/ # all huggingface-cli download mus8tte/lightvc --include 'vocoder/*' --local-dir models/ # vocoder only ``` --- ## vocoder/ — freeC (low-latency neural vocoder) F0-free ISTFT-head neural vocoder (Vocos-style ConvNeXt backbone + complex-STFT head), 27.9 M params, groups=1. - **Quality**: ear-parity with BigVGAN on the target voice (human-ear gate). - **Latency**: causal; synthesis window 5.8 ms; streaming E2E ≈ 39 ms (K=4) incl. mel-analysis lookahead (< 50 ms budget). - **Realtime**: chunked streaming RTF 0.94 @ K=4, single-thread CPU. - **Parity (Rust/Candle vs PyTorch)**: vocoder mel→wave **SNR 88.5 dB**; end-to-end audio→mel→wave resynth **62 dB**. **Files** | file | what | |---|---| | `vocoder/freeC.safetensors` | weights (fp32). Keys: `embed.*`, `blocks.{0..7}.{dw,norm,pw1,pw2}.*`, `norm.*`, `head.*`. | | `vocoder/mel_basis_44k_2048_128.safetensors` | librosa slaney mel filterbank `[128,1025]` (key `mel_basis`) for the input mel extractor. | **Grid / config** - Synthesis: causal, `n_fft = win = 256`, `hop = 128`, `NB = 129`. - Input mel (BigVGAN `mel_spectrogram`): `n_fft = win = 2048`, `hop = 128`, `n_mels = 128`, `sr = 44100`, `fmin = 0`, `fmax = 22050`, Hann, `center=False` + reflect-pad `(n_fft-hop)//2=960`, magnitude `sqrt(·+1e-9)`, `log(clamp(·,1e-5))`. Mel hop = synthesis hop. **Usage** ```bash git clone https://github.com/kjranyone/LightVC && cd LightVC huggingface-cli download mus8tte/lightvc --include 'vocoder/*' --local-dir models_dl/ cargo run -p lightvc-app --release -- resynth \ --input in.wav --weights models_dl/vocoder/freeC.safetensors \ --mel-basis models_dl/vocoder/mel_basis_44k_2048_128.safetensors --output out.wav --k 4 ``` > Note: freeC's "5.8 ms" is the synthesis-window figure; matching-quality streaming needs ~23 ms mel-analysis lookahead (mel trained quasi-centered). A causal-mel retrain restores true ~5.8 ms — see the repo roadmap. **Training**: female multi-speaker corpus, universal vocoder, full-corpus rolling + AMP. No VC-teacher distillation.