Text-to-Speech
Chatterbox
Vietnamese
English
viterbox
viterbox-tts
tts
vietnamese
voice-cloning
audio
speech-synthesis
Instructions to use lukelucky29/viterbox with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Chatterbox
How to use lukelucky29/viterbox with Chatterbox:
# pip install chatterbox-tts import torchaudio as ta from chatterbox.tts import ChatterboxTTS model = ChatterboxTTS.from_pretrained(device="cuda") text = "Ezreal and Jinx teamed up with Ahri, Yasuo, and Teemo to take down the enemy's Nexus in an epic late-game pentakill." wav = model.generate(text) ta.save("test-1.wav", wav, model.sr) # If you want to synthesize with a different voice, specify the audio prompt AUDIO_PROMPT_PATH="YOUR_FILE.wav" wav = model.generate(text, audio_prompt_path=AUDIO_PROMPT_PATH) ta.save("test-2.wav", wav, model.sr) - Notebooks
- Google Colab
- Kaggle
File size: 6,626 Bytes
99713be | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 | ---
license: cc-by-nc-4.0
language:
- vi
- en
library_name: viterbox
tags:
- text-to-speech
- tts
- vietnamese
- voice-cloning
- audio
- speech-synthesis
- chatterbox
datasets:
- vivoice
- phoaudiobook
- dolly-vn/dolly-audio-1000h-vietnamese
pipeline_tag: text-to-speech
base_model: ResembleAI/chatterbox
inference: false
---
<div align="center">
# 🎙️ Viterbox
### Vietnamese Text-to-Speech với Zero-shot Voice Cloning
[](https://creativecommons.org/licenses/by-nc/4.0/)
[](https://www.python.org/downloads/)
[](https://pytorch.org/)
</div>
---
## 📖 Giới thiệu
**Viterbox** là mô hình Text-to-Speech (TTS) tiếng Việt chất lượng cao, được fine-tune từ [Chatterbox](https://github.com/resemble-ai/chatterbox) của Resemble AI.
### ✨ Tính năng chính
- 🇻🇳 **Tiếng Việt tự nhiên**: Phát âm chuẩn, ngữ điệu tự nhiên, hỗ trợ đầy đủ dấu thanh
- 🎯 **Zero-shot Voice Cloning**: Clone giọng nói chỉ với 3-10 giây audio mẫu
- 🌍 **Đa ngôn ngữ**: Hỗ trợ tiếng Việt và 23 ngôn ngữ khác
- ⚡ **Inference nhanh**: Tối ưu GPU với mixed precision (FP16)
- 📝 **Văn bản dài**: Tự động chia câu, ghép audio mượt mà với crossfade
---
### Thông số kỹ thuật
| Thuộc tính | Giá trị |
|------------|---------|
| Sample Rate | 24,000 Hz |
| Audio Format | Mono, 16-bit |
| Vocabulary Size | 2,549 tokens |
| Max Text Length | 2,048 tokens |
| Max Audio Length | ~40 giây |
---
## 📚 Training Data
Model được fine-tune trên **3,000+ giờ** dữ liệu tiếng Việt chất lượng cao:
| Dataset | Mô tả | Thời lượng | Speakers |
|---------|-------|------------|----------|
| **ViVoice** | Dữ liệu giọng đọc đa dạng, nhiều vùng miền | ~1,000h | 500+ |
| **PhoAudiobook** | Sách nói tiếng Việt, giọng chuyên nghiệp | ~1,200h | 100+ |
| **Dolly-Audio** | Dữ liệu nội bộ, đa phong cách | ~800h | 200+ |
### Base Model
- **Chatterbox Multilingual** by Resemble AI
- Pretrained trên 23 ngôn ngữ: Arabic, Chinese, Danish, Dutch, English, Finnish, French, German, Greek, Hebrew, Hindi, Italian, Japanese, Korean, Malay, Norwegian, Polish, Portuguese, Russian, Spanish, Swahili, Swedish, Turkish
- Vocabulary mở rộng thêm tokens cho tiếng Việt
---
## 🚀 Quick Start
### Cài đặt
Hoặc từ source:
```bash
git clone https://github.com/iamdinhthuan/viterbox
cd viterbox
pip install -e .
```
### Sử dụng cơ bản
```python
from viterbox import Viterbox
# Load model (tự động download)
tts = Viterbox.from_pretrained("cuda")
# Generate speech
audio = tts.generate("Xin chào, tôi là Viterbox!")
# Save to file
tts.save_audio(audio, "output.wav")
```
### Voice Cloning
```python
# Clone voice từ audio mẫu (3-10 giây)
audio = tts.generate(
text="Tôi có thể nói bằng giọng của bạn!",
language="vi",
audio_prompt="reference.wav",
exaggeration=0.5,
cfg_weight=0.5,
)
```
### Xử lý văn bản dài
```python
text = """
Việt Nam là một quốc gia nằm ở phía đông bán đảo Đông Dương.
Đất nước có hình chữ S với chiều dài hơn 1600 km.
Thủ đô Hà Nội là trung tâm văn hóa của cả nước.
"""
audio = tts.generate(
text=text,
language="vi",
sentence_pause_ms=500, # Nghỉ 0.5s giữa các câu
)
```
---
## 🎛️ Parameters
| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| `text` | str | required | Văn bản cần đọc |
| `language` | str | `"vi"` | Mã ngôn ngữ (`"vi"` hoặc `"en"`) |
| `audio_prompt` | str/Path | None | Đường dẫn audio mẫu cho voice cloning |
| `exaggeration` | float | 0.5 | Mức độ biểu cảm (0.0-2.0) |
| `cfg_weight` | float | 0.5 | Độ bám sát giọng mẫu (0.0-1.0) |
| `temperature` | float | 0.8 | Độ ngẫu nhiên (0.1-1.0) |
| `top_p` | float | 0.9 | Top-p sampling |
| `repetition_penalty` | float | 1.2 | Penalty cho việc lặp từ |
| `sentence_pause_ms` | int | 500 | Thời gian nghỉ giữa các câu (ms) |
| `crossfade_ms` | int | 50 | Thời gian crossfade khi ghép audio (ms) |
## ⚙️ System Requirements
| Requirement | Minimum | Recommended |
|-------------|---------|-------------|
| Python | 3.10+ | 3.11+ |
| CUDA | 11.8+ | 12.0+ |
| RAM | 8 GB | 16 GB |
| VRAM | 6 GB | 8 GB+ |
---
## ⚠️ Limitations
- Audio mẫu nên sạch, không nhiễu, độ dài 3-10 giây
- Chưa hỗ trợ streaming inference
- Văn bản quá dài (>500 từ) có thể giảm chất lượng
- Một số từ tiếng Anh trong văn bản Việt có thể phát âm chưa chuẩn
---
## 🔒 Ethical Considerations
### Intended Use
- Tạo nội dung audio cho podcast, audiobook, e-learning
- Accessibility tools cho người khiếm thị
- Virtual assistants và chatbots
- Nghiên cứu và phát triển TTS
### Misuse Prevention
- **KHÔNG** sử dụng để tạo deepfake hoặc nội dung lừa đảo
- **KHÔNG** clone giọng nói mà không có sự đồng ý của chủ sở hữu
- **KHÔNG** tạo nội dung vi phạm bản quyền hoặc pháp luật
---
## 📄 License
**CC BY-NC 4.0** (Creative Commons Attribution-NonCommercial 4.0)
- ✅ Sử dụng cho mục đích **phi thương mại**
- ✅ Chia sẻ, sửa đổi với ghi nguồn
- ❌ **KHÔNG** được sử dụng cho mục đích thương mại
Liên hệ thương mại: [contextbox.ai](https://contextbox.ai)
---
## 🙏 Acknowledgements
- **[Resemble AI](https://www.resemble.ai/)** - Chatterbox base model
- **ViVoice, PhoAudiobook** - Vietnamese speech datasets
- **[Dolly VN](https://github.com/dolly-vn)** - Speech Team @ **[ContextBoxAI](https://contextbox.ai)**
---
## 📧 Contact
- **Organization**: [ContextBoxAI](https://contextbox.ai)
- **Team**: [Dolly VN](https://github.com/dolly-vn) - Speech Team
- **GitHub**: [iamdinhthuan/viterbox-tts](https://github.com/iamdinhthuan/viterbox-tts)
- **HuggingFace**: [dolly-vn/viterbox](https://huggingface.co/dolly-vn/viterbox)
---
## 📚 Citation
```bibtex
@misc{viterbox2025,
author = {Dolly VN, ContextBoxAI},
title = {Viterbox: Vietnamese Text-to-Speech with Voice Cloning},
year = {2025},
publisher = {HuggingFace},
url = {https://huggingface.co/dolly-vn/viterbox}
}
``` |