Instructions to use Tharshan/indicf5_hindi-english_code_switch with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- F5-TTS
How to use Tharshan/indicf5_hindi-english_code_switch with F5-TTS:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Indic Code-Switched TTS
Text-to-speech for Indian languages that pronounces embedded English words correctly. ai4bharat/IndicF5 fine-tuned on OpenSLR-104 Hindi-English speech.
Usage
pip install torch torchaudio torchdiffeq x-transformers vocos librosa soundfile transformers
from transformers import AutoModel
import soundfile as sf
model = AutoModel.from_pretrained(
"Tharshan/indicf5_hindi-english_code_switch", trust_remote_code=True)
print(list(model.voices()))
# ['ritu_hinglish', 'ta_hinglish', 'bn', 'gu', 'kn', 'ml', 'mr', 'or', 'pa', 'te']
ref, ref_text = model.voice("ritu_hinglish")
audio, sr = model.generate(
"मैं आज office जा रहा हूँ, morning में एक important meeting है।",
ref_audio=ref, ref_text=ref_text)
sf.write("out.wav", audio, sr)
Your own voice — the transcript must match the clip exactly:
audio, sr = model.generate(text, ref_audio="my.wav", ref_text="exact transcript")
Match the reference voice to your text's language. It affects output quality more than any other setting.
Self-contained: weights, vocabulary, model code and ten reference voices ship in
this repo. No gated downloads, no f5_tts install.
Training
Adapter-style training with a frozen backbone, sized to fit a Kaggle session.
| Base | ai4bharat/IndicF5 (F5-TTS DiT, 0.3B) |
| Data | OpenSLR-104 Hindi-English |
| Method | Adapter / QLoRA, backbone frozen |
| Epochs | 20 |
| Learning rate | 5e-5 |
| Batch size | 19200 frames |
| Gradient accumulation | 4 |
| Warmup steps | 500 |
| Hardware | 2× NVIDIA T4, 32 GB combined (Kaggle) |
| Training time | ~10 hours |
| Vocab | 2546, unchanged |
| Sample rate | 24 kHz |
Batch size is reduced from the 38400-frame default to fit T4 VRAM. A full fine-tune is not feasible in a single Kaggle session on this hardware.
Results
Whisper large-v3 transcripts of generated audio, same reference voice for both models. English words in bold.
Hindi — unchanged. CER 0.000 → 0.000.
| Input | नमस्ते दोस्तों, आज का मौसम बहुत अच्छा है और मैं सोच रहा हूँ कि हम सब मिलकर पार्क में घूमने चलें। |
| Base | नमस्ते दोस्तों! आज का मौसम बहुत अच्छा है और मैं सोच रहा हूँ कि हम सब मिलकर पार्क में घूमने चले। |
| This model | नमस्ते दोस्तों, आज का मौसम बहुत अच्छा है और मैं सोच रहा हूं कि हम सब मिलकर पार्क में घूमने चलें |
Hindi-English — what the model was trained for.
| Input | मैं आज office जा रहा हूँ, morning में एक important meeting है और उसके बाद team के साथ new project discuss करना है। |
| Base | मैं आज अफे जा रहा हूँ ओडे में एक टहौर का निये है और उसके बाद एंग के साथ मिफ्रय इसस करना है |
| This model | मैं आज ओफिस जा रहा हूँ, मॉर्णिंग में एक इंपोर्टेंट मीटिंग है और उसके बाद टीम के साथ नी प्रोजेक्ट डिस्कस करना है |
All six English words come through. Base produces nonsense in their place.
English only — improved but unreliable. CER 0.725 → 0.209.
| Input | The weather is quite pleasant today, so we are planning to visit the park in the evening. |
| Base | Anjay, aur har sone jo isse sehar se hiye |
| This model | Please send today. So we are planning to visit Titey Par. |
Pure English was never the training target — the corpus is English embedded in Hindi. Base reads English text with Indic phonetics and is unusable; this model recovers part of it but drops and garbles words on longer input.
CER is reported only where the scripts match. On code-switched text it ranks the two models backwards: Whisper writes correctly-pronounced English in the local script (morning → मॉर्णिंग), so every correct English word scores as several character errors. Read the transcripts instead.
Other languages
Trained on Hindi-English only. No training data for any language below — yet the English rendering improves in all eight. Input in each case: "I'm going to office today, there's an important meeting in the morning".
| Language | Base IndicF5 | This model |
|---|---|---|
| Tamil | நான் இன்னைக்கு நாப் ஹேவ் போறேன் மூனேல ஒரு எந்தாண்ட நேயை இருக்கு | நான் இன்னைக்கு அவ்வேச் செல்கிறேன் மார்ணிங்கல ஒரு இம்போர்ட்டன்ட் மீட்டி இருக்கிறது |
| Telugu | నేను ఇరోజు నాఫె కి వెళ్లుతున్నాను ఉన్నే లో ఒక ఏహటోన ఈ ఉంది | నేను ఇరోజు ఓఫేస్ కి వేలుతున్నాను మోనింగ్ లో ఒక ఇంపోటంట్ మేటింగ్ ఉంది |
| Bengali | আমি আজাফে জাছি অনি একটা রিফান্ধ হো নেই আছে | আমি আজ ওফিস জাছি মোনিঙে এ একটা ইম্পর |
| Kannada | ನಾನು ಇವತ್ತು ನಭಾಯಗೆ ಹೋಗುತ್ತಿದ್ದೇನೆ ಮೋಣೆನಲ್ಲಿ ಒಂದು ಹೋಣಿ ಇಯಿ ಇದೆ | ನಾನು ಇವತ್ ಓಫಿಸ್ಗೆ ಹೋಗುತ್ತಿದ್ದೇನೆ ಮೋರ್ಣಿಂಗ್ ನಲ್ಲಿ ಒಂದು ಇಂಪಾಟಂಟ್ ಮಿಟಿಂಗ್ ಇದೆ |
| Malayalam | यान इन्न हफले की पोगुनु होरुनेल ओरु टेन्हान मेई उन्ड | ज्यान इन्नो ओफेस लेक् पोगुनु मौर्णिंग ओरु इंपोर्टेंट मीटिंग उन्डो |
| Marathi | मी आज ओफेल आजातोई ओणे मधे एक अहोन इई आहे | मी आज ओफिसला जातोई मौर्णिंग मधेक इंपोर्टेंट मीटिंग आहे |
| Gujarati | હું આજે અફે જઈ રહ્યો છું હોણે માં એક હોર્ચણ એહી છે | હું આજે ઓફેશ જાઈ રહેઓ છું મોર્નિંગમાં એક ઇંપોર્ટન્ટ મેટિંગ છે |
| Punjabi | मੈ ਆਜ ਅਫਜਾਰੇ ਹਾਂ ਹੋਣੇ ਵੀਚ ਇਕ ਥੋਰਚਾ ਨੀਹੀ ਹੈ | मੈ ਆਜ ਆਫੇਸ ਜਾਰੇ ਆਹਾਂ ਮੋਰਣੀਂ ਵੀਚ ਇਕ ਇਮਪੋਟਨਟ ਮੀਟੀਂ ਹੈ |
The capability appears to be script-level rather than tied to Hindi.
Pure-language controls
No English at all — these test whether the fine-tune cost anything.
| Language | Base IndicF5 | This model |
|---|---|---|
| Tamil | வணக்கம் நண்பர்களே! இன்று வானிலை மிகவும் நன்றாக இருக்கிறது. | வணக்கம் நண்பர்களே! இன்று வானிலை மிகவும் நன்றாக இருக்கிறது. |
| Telugu | నమస్కారం మిత్రులార ఇరోజు వాతావరణం చాల బాగుంది | నమస్కారం మిత్రులార ఇరోజు వాతావరనం చాల బాగుంది |
| Bengali | নমস্কার বন্ধুরা, আজ আবোহাওয়া খুপ শুন্দর | নমসকার বন্ধুরা, আজ আভাওয়া খুপ শুন্দর |
Tamil is character-for-character identical. Telugu and Bengali differ by a single character. No catastrophic forgetting, despite a full fine-tune with no replay data.
Reference voices for Odia also ship here, but Whisper has no Odia support, so it isn't measured.
Audio for every row: samples/
Limitations
- Pure English is improved over base but unreliable — words are dropped and garbled on longer input
- Single-pass generation, no text chunking — one or two sentences at a time
- office is correct in Hindi, Kannada and Marathi but garbled in Tamil, Telugu, Gujarati and Punjabi
- Timbre is duller than base — training audio was upsampled from an ASR corpus
- One sentence per language, transcribed by a single ASR model
- Do not clone a voice without the speaker's consent
Credits
ai4bharat/IndicF5 (MIT) · F5-TTS (MIT) · OpenSLR-104
Training details, evaluation scripts and audio demos: GitHub
- Downloads last month
- 93
Model tree for Tharshan/indicf5_hindi-english_code_switch
Base model
ai4bharat/IndicF5