Instructions to use SPRINGLab/SPRING_F5 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use SPRINGLab/SPRING_F5 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-to-speech", model="SPRINGLab/SPRING_F5", trust_remote_code=True)# Load model directly from transformers import SPRING_F5 model = SPRING_F5.from_pretrained("SPRINGLab/SPRING_F5", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Evaluation
Install packages for evaluation:
pip install -e .[eval]
For faster-whisper, for various compatibilities:
pip install ctranslate2==4.5.0if CUDA 12 and cuDNN 9;pip install ctranslate2==4.4.0if CUDA 12 and cuDNN 8;pip install ctranslate2==3.24.0if CUDA 11 and cuDNN 8.
Generating Samples for Evaluation
Prepare Test Datasets
- Seed-TTS testset: Download from seed-tts-eval.
- LibriSpeech test-clean: Download from OpenSLR.
- Unzip the downloaded datasets and place them in the
data/directory. - Our filtered LibriSpeech-PC 4-10s subset:
data/librispeech_pc_test_clean_cross_sentence.lst
Batch Inference for Test Set
To run batch inference for evaluations, execute the following commands:
# if not setup accelerate config yet
accelerate config
# if only perform inference
bash src/f5_tts/eval/eval_infer_batch.sh --infer-only
# if inference and with corresponding evaluation, setup the following tools first
bash src/f5_tts/eval/eval_infer_batch.sh
Objective Evaluation on Generated Results
Download Evaluation Model Checkpoints
- Chinese ASR Model: Paraformer-zh
- English ASR Model: Faster-Whisper
- WavLM Model: Download from Google Drive.
ASR model will be automatically downloaded if
--localnot set for evaluation scripts.
Otherwise, you should update theasr_ckpt_dirpath values ineval_librispeech_test_clean.pyoreval_seedtts_testset.py.WavLM model must be downloaded and your
wavlm_ckpt_dirpath updated ineval_librispeech_test_clean.pyandeval_seedtts_testset.py.
Objective Evaluation Examples
Update the path with your batch-inferenced results, and carry out WER / SIM / UTMOS evaluations:
# Evaluation [WER] for Seed-TTS test [ZH] set
python src/f5_tts/eval/eval_seedtts_testset.py --eval_task wer --lang zh --gen_wav_dir <GEN_WAV_DIR> --gpu_nums 8
# Evaluation [SIM] for LibriSpeech-PC test-clean (cross-sentence)
python src/f5_tts/eval/eval_librispeech_test_clean.py --eval_task sim --gen_wav_dir <GEN_WAV_DIR> --librispeech_test_clean_path <TEST_CLEAN_PATH>
# Evaluation [UTMOS]. --ext: Audio extension
python src/f5_tts/eval/eval_utmos.py --audio_dir <WAV_DIR> --ext wav
Evaluation results can also be found in
_*_results.jsonlfiles saved in<GEN_WAV_DIR>/<WAV_DIR>.