File size: 8,733 Bytes
94ebf1f dcc8074 94ebf1f 4ef546c 94ebf1f 02f99cd e26426a 94ebf1f 81f94e5 94ebf1f 81f94e5 94ebf1f 81f94e5 94ebf1f 81f94e5 94ebf1f e26426a 94ebf1f 4dd4782 94ebf1f 4ef546c 94ebf1f | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 | ---
license: apache-2.0
language:
- hi
- en
tags:
- automatic-speech-recognition
- whisper
- hindi
- hinglish
- code-switching
- speech-recognition
pipeline_tag: automatic-speech-recognition
---
# Tara: Frontier Hindi Transcription Model
> Watch the launch video [here](https://youtu.be/R4-p5JuaBG4). Try via API [here](https://router.trelis.com/models).
**Tara is a frontier automatic speech-recognition model for Hindi and mixed-code (Hinglish)
transcription.** On the AI4Bharat Vistaar Hindi benchmark suite it achieves state-of-the-art
aggregate accuracy, outperforming leading commercial Hindi ASR systems on the 7-benchmark
Vistaar mean, while natively handling Hindi–English code-switched speech through a dedicated **mixed-code
mode** that renders English words in Latin script and Hindi in Devanagari, the way real
Hinglish is written.
## Highlights
- **State-of-the-art Vistaar Hindi aggregate**: 12.06 WER mean over the 7 Vistaar sets, ahead
of Sarvam Saaras-v3 (12.32), with wins on Kathbath, GramVaani, IndicTTS and CommonVoice-hi.
- **Native code-switching**: 8.37 WER on Code-Switch FLEURS (CS-FLEURS) Hindi–English read code-switch via Tara's
mixed-code mode, competitive with the best commercial systems.
- **Robust across domains**: read speech, noisy speech, telephony (GramVaani 21.03 vs Sarvam
23.00), spontaneous conversation (IndicVoices), and accented adult/child speech (HiACC).
- **Bilingual**: retains strong English (6.68 WER CommonVoice-en, 4.55 FLEURS-en).
- **Standard tooling**: loads with 🤗 Transformers exactly like `openai/whisper-large-v3`.
## Usage
```python
import librosa
import torch
from transformers import WhisperProcessor, WhisperForConditionalGeneration
repo = "Trelis/tara"
processor = WhisperProcessor.from_pretrained(repo)
model = WhisperForConditionalGeneration.from_pretrained(
repo, torch_dtype=torch.bfloat16).to("cuda")
tk = processor.tokenizer
hi, en, mc = (tk.convert_tokens_to_ids(t) for t in ("<|hi|>", "<|en|>", "<|mixedcode|>"))
trn, nts = (tk.convert_tokens_to_ids(t) for t in ("<|transcribe|>", "<|notimestamps|>"))
audio_16k, _ = librosa.load("clip.wav", sr=16000, mono=True)
feats = processor(audio_16k, sampling_rate=16000,
return_tensors="pt").input_features.to("cuda", torch.bfloat16)
# Example 1: pure Hindi
out = model.generate(input_features=feats,
forced_decoder_ids=[(1, hi), (2, trn), (3, nts)],
max_new_tokens=444)
print(tk.decode(out[0], skip_special_tokens=True))
# Example 2: Hindi-English mixed-code, inject <|mixedcode|> right after the language token.
# Language auto-detection also works: generate one step unforced and the FIRST generated
# token is the language token; then inject <|mixedcode|> after it and continue.
out = model.generate(input_features=feats,
forced_decoder_ids=[(1, hi), (2, mc), (3, trn), (4, nts)],
max_new_tokens=444)
print(tk.decode(out[0], skip_special_tokens=True))
```
The mixed-code mode (the `<|mixedcode|>` prefix above) conditions generation only: on pure-Hindi
audio it neither degrades accuracy nor forces transliteration; on mixed-code audio it renders
English words in Latin script.
## Evaluation
Evaluation code, the exact text normalizer, and Tara's per-utterance predictions for every
benchmark below are published at [TrelisResearch/tara](https://github.com/TrelisResearch/tara),
so all numbers can be reproduced or re-scored under alternative normalizers.
**Protocol.** All numbers are corpus WER after light text normalization* (Unicode NFC plus
punctuation removal; nukta and all vowel and nasal marks preserved). All systems are scored
on clips ≤ 30 s with identical references. Commercial-system results are measured by us
under the same protocol; they are not vendor-reported figures.
### Vistaar Hindi benchmark (WER ↓)
| Benchmark | **Tara** | Sarvam Saaras-v3 | ElevenLabs Scribe-v2 |
|---|--:|--:|--:|
| [Kathbath](https://huggingface.co/datasets/Trelis/vistaar-hi-kathbath-test) (clean read) | **9.34** | 9.71 | 9.60 |
| [Kathbath-hard](https://huggingface.co/datasets/Trelis/vistaar-hi-kathbath_noisy-test) (noisy) | 10.82 | **10.55** | 11.11 |
| [MUCS](https://huggingface.co/datasets/Trelis/vistaar-hi-mucs-test) | 10.79 | **9.69** | 10.93 |
| [GramVaani](https://huggingface.co/datasets/Trelis/vistaar-hi-gramvaani-test) (telephony) | **21.03** | 23.00 | 26.94 |
| [IndicTTS](https://huggingface.co/datasets/Trelis/vistaar-hi-indictts-test) | **9.46** | 10.38 | 13.17 |
| [CommonVoice-hi](https://huggingface.co/datasets/Trelis/vistaar-hi-commonvoice-test) | **12.51** | 12.88 | 13.44 |
| [FLEURS-hi](https://huggingface.co/datasets/Trelis/vistaar-hi-fleurs-test) | 10.47 | **10.05** | 11.33 |
| **Mean (7 Vistaar sets)** | **12.06** | 12.32 | 13.79 |
| [IndicVoices-500](https://huggingface.co/datasets/Trelis/indicvoices-500-hi-eval) (spontaneous, non-Vistaar) | 16.51 | **15.29** | 27.46 |
IndicVoices-500 is a 500-sample spontaneous-speech control from the IndicVoices validation
split; it is not part of the Vistaar mean.
### Code-switching (Hinglish) benchmarks (WER ↓)
Tara and Sarvam are measured in their code-mixed modes.
| Benchmark | **Tara** | Sarvam Saaras-v3 | ElevenLabs Scribe-v2 |
|---|--:|--:|--:|
| [CoSHE-500](https://huggingface.co/datasets/Trelis/CoSHE-500) (conversational CS) | 14.41 | **11.25** | 12.40 |
| [Code-Switch FLEURS hi-en](https://huggingface.co/datasets/Trelis/cs-fleurs-hineng-read-test) (read CS) | 8.37 | 16.47 | **7.57** |
| [Hi-accent adult (HiACC)](https://huggingface.co/datasets/Trelis/hiacc-adult-test-eval) | 12.93 | 13.16 | **12.87** |
| [Hi-accent child (HiACC)](https://huggingface.co/datasets/Trelis/hiacc-child-test-eval) | 10.69 | **10.10** | 11.66 |
### English (WER ↓)
Scored with the standard Whisper English normalizer.
| Benchmark | **Tara** | Sarvam | Scribe-v2 |
|---|--:|--:|--:|
| [CommonVoice-en](https://huggingface.co/datasets/Trelis/cv-en-scripted-test-500) | 6.68 | 8.68 | **5.28** |
| [FLEURS-en](https://huggingface.co/datasets/Trelis/fleurs-en-test) | 4.55 | 4.36 | **2.93** |
\* Normalization: `unicodedata.normalize("NFC")`, lowercasing, then removal of punctuation
and symbols (`। , . ? ! " : ; - – — “ ” ( ) [ ] < > / ~ % ₹ $ …`), invisible formatting
characters (zero-width joiner/space) and the Unicode replacement character; apostrophes are
kept. The ≤30 s rule excludes 2 clips on GramVaani, 2 on IndicTTS and 1 on FLEURS-hi; no
other set has any. Measurement error is small: re-runs across hardware and precision agree
to within 0.1 WER. There is also slight noise in the reference labels (for example
inconsistent nukta spelling: both हज़ार and हजार appear as references within GramVaani, and
both ज़्यादा and ज्यादा within MUCS), but this should not affect the numbers by much.
## Limitations
- **Mode selection**: peak accuracy comes from picking the mode per clip (Hindi, mixed-code, or
English). When the language mix is unknown, **the Hindi mixed-code mode is a safe default**: on
pure-Hindi audio it produces pure-Hindi transcripts with no measured accuracy loss, and on mixed
audio it handles the code-switching. Automatic language detection is also supported: let the
model generate the language token and inject mixed-code after it (see Usage).
- **Clip length**: evaluated on clips ≤ 30 s; longer audio should be chunked (standard Whisper
practice).
- Hindi–English only; other Indic languages are out of scope for this release.
## Intended use
Transcription of Hindi and Hindi–English code-switched speech: voice assistants, contact-center
analytics, media captioning, and speech data pipelines.
## Model details
- **Architecture**: Whisper large-v3 (encoder–decoder, 1.55B params) + mixed-code mode
- **Languages**: Hindi (hi), English (en), Hindi–English code-switch
- **Sample rate**: 16 kHz input
- **I/O**: ≤30 s audio per window → text
- **License**: Apache 2.0
## Attribution
We thank [Gram Vaani](https://gramvaani.org) for permission to use the
[Gram Vaani ASR Challenge 2022 Corpus](https://www.openslr.org/118/) in training Tara. Gram Vaani
builds community-anchored voice based engagement platforms ('Mobile Vaani' clubs) that give underserved and
marginalised communities a channel to access information and express themselves.
## License
This model is released under the **Apache License 2.0**.
## Citation
If you use Tara in your work, please cite:
```bibtex
@misc{trelis2026tara,
title = {Tara: Frontier Hindi Transcription Model},
author = {{Trelis Research}},
year = {2026},
url = {https://huggingface.co/Trelis/tara}
}
```
|