EditVoice โ€” Official Model Weights

This is the official model weights repository for EditVoice: Variable-Length Non-Autoregressive Zero-Shot TTS and Speech Editing with Edit Flows. It provides the inference checkpoints and Edit Flow initial distribution for zero-shot TTS and speech editing in noisy environments. The official code repository contains the inference implementation, and audio examples are available in the EditVoice demo.

Files and Provenance

File Purpose Source and license
checkpoints/editvoice_lm.pt Edit Flow semantic model for TTS and speech editing EditVoice; Apache-2.0
checkpoints/kimi-flow.pt Token-to-mel Flow Matching model fine-tuned for speech editing in noisy environments; paired with checkpoints/kimi-bigvgan/ EditVoice fine-tuning from CosyVoice2 flow.pt; Apache-2.0
data/empirical.pkl 6,561-class initial distribution $p_0$ used by Edit Flows EditVoice training-data statistics; Apache-2.0
checkpoints/flow.pt Token-to-mel Flow Matching model for zero-shot TTS; paired with checkpoints/hift.pt CosyVoice2-0.5B; Apache-2.0
checkpoints/hift.pt Vocoder for zero-shot TTS; paired with checkpoints/flow.pt CosyVoice2-0.5B; Apache-2.0
checkpoints/speech_tokenizer_v2.onnx Speech tokenizer CosyVoice2-0.5B; Apache-2.0
checkpoints/campplus.onnx Speaker embedding extractor CosyVoice2-0.5B; Apache-2.0
checkpoints/kimi-bigvgan/config.json, checkpoints/kimi-bigvgan/model.pt Vocoder for speech editing in noisy environments; paired with checkpoints/kimi-flow.pt vocoder/ from Kimi-Audio-7B-Instruct; MIT

The ttsfrd wheel and its English resources are not included; obtain them separately from an authorized source.

Download and Use

With the Hugging Face CLI installed, download the files and verify their checksums:

hf download DOOD02/EditVoice --local-dir model_assets
(cd model_assets && sha256sum -c SHA256SUMS)

For editvoice-tts, pass model_assets/checkpoints/editvoice_lm.pt, flow.pt, hift.pt, speech_tokenizer_v2.onnx, campplus.onnx, and model_assets/data/empirical.pkl to their corresponding checkpoint and resource arguments. For editvoice-edit, use the same semantic model, tokenizer, speaker model, and distribution, together with model_assets/checkpoints/kimi-flow.pt and model_assets/checkpoints/kimi-bigvgan/. See the EditVoice code repository README for complete commands and input formats.

License

This repository contains files under Apache-2.0 and MIT; no single license applies to every file. See LICENSE for the file-by-file license mapping and licenses/ for the license texts. Third-party weights retain their original authorship and licenses.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Paper for DOOD02/EditVoice