Indic Code-Switched TTS

Text-to-speech for Indian languages that pronounces embedded English words correctly. ai4bharat/IndicF5 fine-tuned on OpenSLR-104 Hindi-English speech.

Usage

pip install torch torchaudio torchdiffeq x-transformers vocos librosa soundfile transformers
from transformers import AutoModel
import soundfile as sf

model = AutoModel.from_pretrained(
    "Tharshan/indicf5_hindi-english_code_switch", trust_remote_code=True)

print(list(model.voices()))
# ['ritu_hinglish', 'ta_hinglish', 'bn', 'gu', 'kn', 'ml', 'mr', 'or', 'pa', 'te']

ref, ref_text = model.voice("ritu_hinglish")
audio, sr = model.generate(
    "मैं आज office जा रहा हूँ, morning में एक important meeting है।",
    ref_audio=ref, ref_text=ref_text)
sf.write("out.wav", audio, sr)

Your own voice — the transcript must match the clip exactly:

audio, sr = model.generate(text, ref_audio="my.wav", ref_text="exact transcript")

Match the reference voice to your text's language. It affects output quality more than any other setting.

Self-contained: weights, vocabulary, model code and ten reference voices ship in this repo. No gated downloads, no f5_tts install.

Training

Adapter-style training with a frozen backbone, sized to fit a Kaggle session.

Base ai4bharat/IndicF5 (F5-TTS DiT, 0.3B)
Data OpenSLR-104 Hindi-English
Method Adapter / QLoRA, backbone frozen
Epochs 20
Learning rate 5e-5
Batch size 19200 frames
Gradient accumulation 4
Warmup steps 500
Hardware 2× NVIDIA T4, 32 GB combined (Kaggle)
Training time ~10 hours
Vocab 2546, unchanged
Sample rate 24 kHz

Batch size is reduced from the 38400-frame default to fit T4 VRAM. A full fine-tune is not feasible in a single Kaggle session on this hardware.

Results

Whisper large-v3 transcripts of generated audio, same reference voice for both models. English words in bold.

Hindi — unchanged. CER 0.000 → 0.000.

Input नमस्ते दोस्तों, आज का मौसम बहुत अच्छा है और मैं सोच रहा हूँ कि हम सब मिलकर पार्क में घूमने चलें।
Base नमस्ते दोस्तों! आज का मौसम बहुत अच्छा है और मैं सोच रहा हूँ कि हम सब मिलकर पार्क में घूमने चले।
This model नमस्ते दोस्तों, आज का मौसम बहुत अच्छा है और मैं सोच रहा हूं कि हम सब मिलकर पार्क में घूमने चलें

Hindi-English — what the model was trained for.

Input मैं आज office जा रहा हूँ, morning में एक important meeting है और उसके बाद team के साथ new project discuss करना है।
Base मैं आज अफे जा रहा हूँ ओडे में एक टहौर का निये है और उसके बाद एंग के साथ मिफ्रय इसस करना है
This model मैं आज ओफिस जा रहा हूँ, मॉर्णिंग में एक इंपोर्टेंट मीटिंग है और उसके बाद टीम के साथ नी प्रोजेक्ट डिस्कस करना है

All six English words come through. Base produces nonsense in their place.

English only — improved but unreliable. CER 0.725 → 0.209.

Input The weather is quite pleasant today, so we are planning to visit the park in the evening.
Base Anjay, aur har sone jo isse sehar se hiye
This model Please send today. So we are planning to visit Titey Par.

Pure English was never the training target — the corpus is English embedded in Hindi. Base reads English text with Indic phonetics and is unusable; this model recovers part of it but drops and garbles words on longer input.

CER is reported only where the scripts match. On code-switched text it ranks the two models backwards: Whisper writes correctly-pronounced English in the local script (morning → मॉर्णिंग), so every correct English word scores as several character errors. Read the transcripts instead.

Other languages

Trained on Hindi-English only. No training data for any language below — yet the English rendering improves in all eight. Input in each case: "I'm going to office today, there's an important meeting in the morning".

Language Base IndicF5 This model
Tamil நான் இன்னைக்கு நாப் ஹேவ் போறேன் மூனேல ஒரு எந்தாண்ட நேயை இருக்கு நான் இன்னைக்கு அவ்வேச் செல்கிறேன் மார்ணிங்கல ஒரு இம்போர்ட்டன்ட் மீட்டி இருக்கிறது
Telugu నేను ఇరోజు నాఫె కి వెళ్లుతున్నాను ఉన్నే లో ఒక ఏహటోన ఈ ఉంది నేను ఇరోజు ఓఫేస్ కి వేలుతున్నాను మోనింగ్ లో ఒక ఇంపోటంట్ మేటింగ్ ఉంది
Bengali আমি আজাফে জাছি অনি একটা রিফান্ধ হো নেই আছে আমি আজ ওফিস জাছি মোনিঙে এ একটা ইম্পর
Kannada ನಾನು ಇವತ್ತು ನಭಾಯಗೆ ಹೋಗುತ್ತಿದ್ದೇನೆ ಮೋಣೆನಲ್ಲಿ ಒಂದು ಹೋಣಿ ಇಯಿ ಇದೆ ನಾನು ಇವತ್ ಓಫಿಸ್ಗೆ ಹೋಗುತ್ತಿದ್ದೇನೆ ಮೋರ್ಣಿಂಗ್ ನಲ್ಲಿ ಒಂದು ಇಂಪಾಟಂಟ್ ಮಿಟಿಂಗ್ ಇದೆ
Malayalam यान इन्न हफले की पोगुनु होरुनेल ओरु टेन्हान मेई उन्ड ज्यान इन्नो ओफेस लेक् पोगुनु मौर्णिंग ओरु इंपोर्टेंट मीटिंग उन्डो
Marathi मी आज ओफेल आजातोई ओणे मधे एक अहोन इई आहे मी आज ओफिसला जातोई मौर्णिंग मधेक इंपोर्टेंट मीटिंग आहे
Gujarati હું આજે અફે જઈ રહ્યો છું હોણે માં એક હોર્ચણ એહી છે હું આજે ઓફેશ જાઈ રહેઓ છું મોર્નિંગમાં એક ઇંપોર્ટન્ટ મેટિંગ છે
Punjabi मੈ ਆਜ ਅਫਜਾਰੇ ਹਾਂ ਹੋਣੇ ਵੀਚ ਇਕ ਥੋਰਚਾ ਨੀਹੀ ਹੈ मੈ ਆਜ ਆਫੇਸ ਜਾਰੇ ਆਹਾਂ ਮੋਰਣੀਂ ਵੀਚ ਇਕ ਇਮਪੋਟਨਟ ਮੀਟੀਂ ਹੈ

The capability appears to be script-level rather than tied to Hindi.

Pure-language controls

No English at all — these test whether the fine-tune cost anything.

Language Base IndicF5 This model
Tamil வணக்கம் நண்பர்களே! இன்று வானிலை மிகவும் நன்றாக இருக்கிறது. வணக்கம் நண்பர்களே! இன்று வானிலை மிகவும் நன்றாக இருக்கிறது.
Telugu నమస్కారం మిత్రులార ఇరోజు వాతావరణం చాల బాగుంది నమస్కారం మిత్రులార ఇరోజు వాతావరనం చాల బాగుంది
Bengali নমস্কার বন্ধুরা, আজ আবোহাওয়া খুপ শুন্দর নমসকার বন্ধুরা, আজ আভাওয়া খুপ শুন্দর

Tamil is character-for-character identical. Telugu and Bengali differ by a single character. No catastrophic forgetting, despite a full fine-tune with no replay data.

Reference voices for Odia also ship here, but Whisper has no Odia support, so it isn't measured.

Audio for every row: samples/

Limitations

  • Pure English is improved over base but unreliable — words are dropped and garbled on longer input
  • Single-pass generation, no text chunking — one or two sentences at a time
  • office is correct in Hindi, Kannada and Marathi but garbled in Tamil, Telugu, Gujarati and Punjabi
  • Timbre is duller than base — training audio was upsampled from an ASR corpus
  • One sentence per language, transcribed by a single ASR model
  • Do not clone a voice without the speaker's consent

Credits

ai4bharat/IndicF5 (MIT) · F5-TTS (MIT) · OpenSLR-104

Training details, evaluation scripts and audio demos: GitHub

Downloads last month
93
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Tharshan/indicf5_hindi-english_code_switch

Finetuned
(17)
this model

Space using Tharshan/indicf5_hindi-english_code_switch 1