1 📄
Input
Text & Ref Audio
2 🔄
Preprocessing
Soundfile & Whisper
3 🧠
F5-TTS (DiT)
Diffusion Transformer
4 🎼
Vocos
Mel -> Waveform
5 🔊
Output
wav file
🧪
Data Prep
🏋️
Training
📊
Evaluation
Use Cases
📚
Glossary
🎛️
Parameters
🧬
F5-TTS Diagram
🧩
Other Models
Requirements
🗂️
Data Layout
🐳
Docker Env
⚙️
ARM64 Fixes