| --- |
| license: mit |
| library_name: candle |
| tags: |
| - voice-conversion |
| - vocoder |
| - real-time |
| - low-latency |
| - rust |
| - candle |
| - asmr |
| language: |
| - ja |
| pipeline_tag: audio-to-audio |
| --- |
| |
| # LightVC — model suite |
|
|
| Own-weights model suite for real-time Japanese voice conversion (ASMR / バ美声). |
| Inference is **pure Rust / [Candle](https://github.com/huggingface/candle)** (no Python runtime), |
| target **E2E < 50 ms, CPU, causal**. Code: <https://github.com/kjranyone/LightVC>. License **MIT**. |
|
|
| All weights are **100% own** (no distillation from other voice-conversion models; targets are real audio only). |
|
|
| ## Components |
|
|
| | component | folder | status | notes | |
| |---|---|---|---| |
| | **vocoder (freeC)** | `vocoder/` | ✅ released | low-latency neural vocoder, ear-parity with BigVGAN | |
| | content encoder (G-enc) | `content-encoder/` | ⬜ planned | causal HuBERT-distilled, leakage-suppressed | |
| | VC generator (AdaIN + CIPT) | `vc/` | ⬜ planned | mel generator, multi-reference moe-factor compositor | |
| | prosody / s_art | `prosody/` | ⬜ planned | articulation clone, prosody policy | |
| | voices | `voices/` | ⬜ planned | target / reference embeddings for the compositor | |
| |
| Download the whole suite or one component: |
| |
| ```bash |
| huggingface-cli download mus8tte/lightvc --local-dir models/ # all |
| huggingface-cli download mus8tte/lightvc --include 'vocoder/*' --local-dir models/ # vocoder only |
| ``` |
| |
| --- |
| |
| ## vocoder/ — freeC (low-latency neural vocoder) |
| |
| F0-free ISTFT-head neural vocoder (Vocos-style ConvNeXt backbone + complex-STFT head), 27.9 M params, groups=1. |
| |
| - **Quality**: ear-parity with BigVGAN on the target voice (human-ear gate). |
| - **Latency**: causal; synthesis window 5.8 ms; streaming E2E ≈ 39 ms (K=4) incl. mel-analysis lookahead (< 50 ms budget). |
| - **Realtime**: chunked streaming RTF 0.94 @ K=4, single-thread CPU. |
| - **Parity (Rust/Candle vs PyTorch)**: vocoder mel→wave **SNR 88.5 dB**; end-to-end audio→mel→wave resynth **62 dB**. |
| |
| **Files** |
| |
| | file | what | |
| |---|---| |
| | `vocoder/freeC.safetensors` | weights (fp32). Keys: `embed.*`, `blocks.{0..7}.{dw,norm,pw1,pw2}.*`, `norm.*`, `head.*`. | |
| | `vocoder/mel_basis_44k_2048_128.safetensors` | librosa slaney mel filterbank `[128,1025]` (key `mel_basis`) for the input mel extractor. | |
|
|
| **Grid / config** |
|
|
| - Synthesis: causal, `n_fft = win = 256`, `hop = 128`, `NB = 129`. |
| - Input mel (BigVGAN `mel_spectrogram`): `n_fft = win = 2048`, `hop = 128`, `n_mels = 128`, `sr = 44100`, `fmin = 0`, `fmax = 22050`, Hann, `center=False` + reflect-pad `(n_fft-hop)//2=960`, magnitude `sqrt(·+1e-9)`, `log(clamp(·,1e-5))`. Mel hop = synthesis hop. |
|
|
| **Usage** |
|
|
| ```bash |
| git clone https://github.com/kjranyone/LightVC && cd LightVC |
| huggingface-cli download mus8tte/lightvc --include 'vocoder/*' --local-dir models_dl/ |
| cargo run -p lightvc-app --release -- resynth \ |
| --input in.wav --weights models_dl/vocoder/freeC.safetensors \ |
| --mel-basis models_dl/vocoder/mel_basis_44k_2048_128.safetensors --output out.wav --k 4 |
| ``` |
|
|
| > Note: freeC's "5.8 ms" is the synthesis-window figure; matching-quality streaming needs ~23 ms mel-analysis lookahead (mel trained quasi-centered). A causal-mel retrain restores true ~5.8 ms — see the repo roadmap. |
|
|
| **Training**: female multi-speaker corpus, universal vocoder, full-corpus rolling + AMP. No VC-teacher distillation. |
|
|