RIMA-TTS-v1 / README.md
Praha-Labs's picture
Upload RIMA-TTS v1 multilingual Indic adapters
a0cac50 verified
|
Raw
History Blame Contribute Delete
3.15 kB
---
license: other
base_model: ResembleAI/chatterbox
library_name: peft
tags:
- text-to-speech
- tts
- chatterbox
- lora
- indic
- multilingual
- hindi
- tamil
- telugu
- malayalam
- kannada
- bengali
- marathi
- gujarati
- punjabi
- urdu
- odia
- assamese
---
# RIMA-TTS v1
RIMA-TTS v1 is a multilingual Indic LoRA adapter for Chatterbox TTS.
This is not a standalone merged model. It is a PEFT/LoRA adapter that must be used with the Chatterbox TTS base model and the included Indic tokenizer.
## Recommended Adapter
Recommended checkpoint: `adapters/checkpoint-12000`.
The final adapter is also included at `adapters/final`, but checkpoint `12000` is recommended based on listening tests.
## Supported Languages
| Code | Language |
|---|---|
| hi | Hindi |
| ta | Tamil |
| te | Telugu |
| ml | Malayalam |
| kn | Kannada |
| bn | Bengali |
| mr | Marathi |
| gu | Gujarati |
| pa | Punjabi |
| ur | Urdu |
| or | Odia |
| as | Assamese |
## Included Adapters
- `adapters/checkpoint-8000`
- `adapters/checkpoint-9000`
- `adapters/checkpoint-10000`
- `adapters/checkpoint-11000`
- `adapters/checkpoint-12000`
- `adapters/checkpoint-12068`
- `adapters/final`
Each adapter folder contains `adapter_config.json` and `adapter_model.safetensors`.
## Training Data
- Dataset: `ai4bharat/Rasa`
- Languages: 12 Indic languages
- Target duration: 20 hours per language
- Total target duration: about 240 hours
- Training rows: 144,812
## Training Configuration
| Setting | Value |
|---|---|
| Base model | Chatterbox TTS |
| Training type | LoRA adapter finetune |
| Epochs | 2 |
| Steps | 12,068 |
| Batch size | 24 |
| Gradient accumulation | 1 |
| Learning rate | 6e-5 |
| LoRA rank | 128 |
| LoRA alpha | 256 |
| Tokenizer vocab size | 3,358 |
| Trainable params | 97,480,704 |
| Total params | 635,321,344 |
| Trainable % | 15.34% |
## Training Result
| Metric | Value |
|---|---|
| Final train loss | 4.0169 |
| Final logged loss | 4.5739 |
| Runtime | 5h 33m 44s |
| Samples/sec | 14.463 |
| Steps/sec | 0.603 |
## Evaluation
Fixed multilingual eval samples were generated every 1000 steps across all 12 languages.
Manual listening tests showed strong early performance by checkpoint 2000 for Hindi, Tamil, and Malayalam. Checkpoint 12000 is currently recommended after later listening comparison.
## Files
- `tokenizer/tokenizer_indic_12lang.json`
- `configs/strong_run_config.py`
- `configs/train_multilingual.py`
- `training/strong_run_train.log`
- `training/dataset_summary.json`
## Usage Notes
This adapter requires:
1. Chatterbox TTS base model files
2. `tokenizer/tokenizer_indic_12lang.json`
3. One adapter folder from `adapters/`
Use `adapters/checkpoint-12000` unless comparing checkpoints.
## Limitations
- This is an adapter only, not a full merged model.
- Quality may vary by language.
- Voice cloning quality depends on the reference audio.
- Formal benchmark metrics such as MOS, WER, CER, or speaker similarity are not included yet.
- Evaluate carefully before production use.
## Attribution
- Base model: Chatterbox TTS
- Training data: ai4bharat/Rasa
- Finetuning: RIMA-TTS v1 multilingual Indic adapter