LightVC — model suite

Own-weights model suite for real-time Japanese voice conversion (ASMR / バ美声). Inference is pure Rust / Candle (no Python runtime), target **E2E < 50 ms, CPU, causal**. Code: https://github.com/kjranyone/LightVC. License MIT.

All weights are 100% own (no distillation from other voice-conversion models; targets are real audio only).

Components

component folder status notes
vocoder (freeC) vocoder/ ✅ released low-latency neural vocoder, ear-parity with BigVGAN
content encoder (G-enc) content-encoder/ ⬜ planned causal HuBERT-distilled, leakage-suppressed
VC generator (AdaIN + CIPT) vc/ ⬜ planned mel generator, multi-reference moe-factor compositor
prosody / s_art prosody/ ⬜ planned articulation clone, prosody policy
voices voices/ ⬜ planned target / reference embeddings for the compositor

Download the whole suite or one component:

huggingface-cli download mus8tte/lightvc --local-dir models/                 # all
huggingface-cli download mus8tte/lightvc --include 'vocoder/*' --local-dir models/   # vocoder only

vocoder/ — freeC (low-latency neural vocoder)

F0-free ISTFT-head neural vocoder (Vocos-style ConvNeXt backbone + complex-STFT head), 27.9 M params, groups=1.

  • Quality: ear-parity with BigVGAN on the target voice (human-ear gate).
  • Latency: causal; synthesis window 5.8 ms; streaming E2E ≈ 39 ms (K=4) incl. mel-analysis lookahead (< 50 ms budget).
  • Realtime: chunked streaming RTF 0.94 @ K=4, single-thread CPU.
  • Parity (Rust/Candle vs PyTorch): vocoder mel→wave SNR 88.5 dB; end-to-end audio→mel→wave resynth 62 dB.

Files

file what
vocoder/freeC.safetensors weights (fp32). Keys: embed.*, blocks.{0..7}.{dw,norm,pw1,pw2}.*, norm.*, head.*.
vocoder/mel_basis_44k_2048_128.safetensors librosa slaney mel filterbank [128,1025] (key mel_basis) for the input mel extractor.

Grid / config

  • Synthesis: causal, n_fft = win = 256, hop = 128, NB = 129.
  • Input mel (BigVGAN mel_spectrogram): n_fft = win = 2048, hop = 128, n_mels = 128, sr = 44100, fmin = 0, fmax = 22050, Hann, center=False + reflect-pad (n_fft-hop)//2=960, magnitude sqrt(·+1e-9), log(clamp(·,1e-5)). Mel hop = synthesis hop.

Usage

git clone https://github.com/kjranyone/LightVC && cd LightVC
huggingface-cli download mus8tte/lightvc --include 'vocoder/*' --local-dir models_dl/
cargo run -p lightvc-app --release -- resynth \
  --input in.wav --weights models_dl/vocoder/freeC.safetensors \
  --mel-basis models_dl/vocoder/mel_basis_44k_2048_128.safetensors --output out.wav --k 4

Note: freeC's "5.8 ms" is the synthesis-window figure; matching-quality streaming needs ~23 ms mel-analysis lookahead (mel trained quasi-centered). A causal-mel retrain restores true ~5.8 ms — see the repo roadmap.

Training: female multi-speaker corpus, universal vocoder, full-corpus rolling + AMP. No VC-teacher distillation.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support