Indic ASR Datasets
Open speech datasets for training Indic ASR models across all 22 scheduled Indian languages.
Viewer • Updated • 6.38M • 18.4k • 85Note 6,380 hrs, 22 langs, CC-BY-4.0 — primary ASR training
ai4bharat/Rasa
Viewer • Updated • 1.28M • 9.3k • 47Note 1,280 hrs — conversational/expressive speech
ai4bharat/Svarah
Viewer • Updated • 6.66k • 795 • 31Note 6.66K rows, 22 langs — evaluation benchmark
ai4bharat/SpeechArenaBench
Viewer • Updated • 125k • 270 • 7Note 125K rows, 22 langs — noisy/code-switched eval
ai4bharat/Rural_Women_ASR_v2
Viewer • Updated • 64.4k • 1.19k • 8Note 64 hrs, 6 langs — far-field/noisy real-world speech
ai4bharat/Rural_Women_Bhojpuri
Viewer • Updated • 78.8k • 308 • 6Note 79 hrs, bhojpuri — low-resource noisy speech
ai4bharat/Shrutilipi
Viewer • Updated • 2.28M • 1.41k • 16Note 2,280 hrs, 12 langs, CC-BY-4.0 — read speech
ai4bharat/Kathbath
Viewer • Updated • 806k • 2.38k • 21Note ~1,000 hrs — speaker variation TTS/ASR
ai4bharat/indicvoices_r
Viewer • Updated • 665k • 6.6k • 32Note 1,704 hrs, 22 langs — TTS-quality augmented speech
ai4bharat/IndicVoices-ST
Viewer • Updated • 1.38M • 264 • 21Note Speech translation variant of IndicVoices
ai4bharat/indicvoices-cleaned
Preview • Updated • 32Note Cleaned subset of IndicVoices
ai4bharat/Lahaja
Viewer • Updated • 6.15k • 198 • 4Note Dialect/accent variation speech data
ai4bharat/Spoken-Tutorial
Viewer • Updated • 248k • 66 • 2Note Educational spoken tutorials in Indic languages
ai4bharat/Mann-ki-Baat
Viewer • Updated • 242k • 99 • 5Note Prime Minister's speech data for Indic languages
ai4bharat/sangraha
Viewer • Updated • 268M • 11.4k • 79Note Compilation of Indic speech resources
mozilla-foundation/common_voice_17_0
Updated • 4.3k • 38Note Crowdsourced read speech — 10 Indic languages, CC-0
iitm-ddp/iiith-indic-speech
Viewer • Updated • 9.64k • 28 • 3Note IIIT-H LTRC collected Indic speech across multiple languages
equal-ai/ultravox-indic-speech
Viewer • Updated • 5.17M • 11 • 1Note 4,200 hrs English+Hindi large-scale multilingual speech
ssws3/Indic-Speech-Datasets
Viewer • Updated • 295k • 461Note Multi-domain Hindi speech combining STT, TTS, and mixed sources
parambharat/telugu_asr_corpus
Updated • 19Note 360 hrs Telugu ASR corpus with de-duplicated transcripts
zsy12345/telugu-asr
Viewer • Updated • 168k • 133 • 2Note Telugu ASR speech dataset
deepdml/iisc-mile-tamil-asr
Viewer • Updated • 89.4k • 196Note IISc-MILE Tamil ASR corpus — large-scale transcribed Tamil speech
parambharat/tamil_asr_corpus
Updated • 36 • 7Note 1,000 hrs Tamil ASR corpus with de-duplicated transcripts
ketav/parakeet-hindi-asr
Viewer • Updated • 373k • 45Note Hindi-English bilingual ASR for Parakeet fine-tuning
ketav/hindi-youtube-asr-transcripts
Updated • 30Note Auto-generated YouTube transcripts from 21 Hindi channels
collabora/hindi-asr-wds
Viewer • Updated • 1.28M • 746Note Hindi ASR dataset in WebDataset format
SkunkWorkLabs/hindi-asr-benchmark
Viewer • Updated • 10k • 48 • 2Note Hindi ASR benchmark evaluating STT models
shujaAK/hindi-dairy-asr-clean
Viewer • Updated • 2.09k • 22Note Clean Hindi dairy-domain ASR dataset
shujaAK/hindi-hinglish-business-asr
Viewer • Updated • 200 • 22Note Hindi-Hinglish business-domain ASR
SKNahin/open-large-bengali-asr-data
Viewer • Updated • 3.73M • 490 • 12Note 5,000 hrs Bengali ASR data aggregated from public sources
parambharat/bengali_asr_corpus
Updated • 24Note 500 hrs Bengali ASR corpus with de-duplicated transcripts
IntisarUddin/Bengali_Long_form_ASR
Updated • 48Note Bengali long-form ASR dataset
maddi99/bengali_english_cs_asr
Viewer • Updated • 2 • 6Note Bengali-English code-switched ASR dataset
JKA-NLP/unified-kannada-asr-1.0
Viewer • Updated • 110k • 320 • 1Note Unified Kannada ASR dataset v1.0 from multiple sources
arpit-tiwari/iisc-mile-kannada-asr-corpus
Viewer • Updated • 106k • 100Note IISc-MILE Kannada ASR corpus
parambharat/kannada_asr_corpus
Updated • 10Note 360 hrs Kannada ASR corpus with de-duplicated transcripts
aoxo/asr_malayalam
Preview • Updated • 228 • 1Note Malayalam ASR dataset
asr-malayalam/combined_malayalam
Viewer • Updated • 89.4k • 190Note Combined Malayalam ASR dataset
sajilck/malayalam-asr-corpus
Viewer • Updated • 96.6k • 74Note Multi-corpus Malayalam speech dataset aggregated from 5 public sources
parambharat/malayalam_asr_corpus
Updated • 21 • 3Note Malayalam ASR corpus
kdcyberdude/Punjabi_ASR_datasets
Viewer • Updated • 130k • 669 • 7Note Punjabi ASR datasets collection (687 downloads)
aaparajit02/punjabi-asr
Viewer • Updated • 39.2k • 182 • 3Note Punjabi ASR dataset filtered from Shrutilipi
MatrixSpeechAI/All_Marathi_ASR
Viewer • Updated • 24.5k • 42 • 1Note Marathi ASR dataset (f)
rootflo/marathi-asr-data
Viewer • Updated • 10.3k • 24Note Marathi ASR data
transitionGap/ASR_Marathi_Sentences
Viewer • Updated • 8.96k • 16Note Marathi sentence-level ASR dataset
haideraqeeb/gujarati-asr-16kHz
Viewer • Updated • 81k • 52 • 1Note Gujarati ASR dataset at 16kHz
XKaab/ASR-gujarati_100hrs
Viewer • Updated • 51.3k • 26Note 100 hrs Gujarati ASR dataset
Minutor/odia_asr_collection
Viewer • Updated • 1.1k • 33Note Aggregated Odia audio datasets for ASR research
Mohan-diffuser/odia-english-ASR
Viewer • Updated • 2.36k • 35Note Odia-English ASR dataset
mahendraphd/Indic_Hindi-English_Parallel_Speech
Updated • 19 • 2Note Hindi-English parallel speech for code-switched ASR
shunyalabs/synthetic-speech-indic
Viewer • Updated • 406k • 10 • 1Note Synthetic Indic speech data
bc7ec356/synthetic-speech-indic
Viewer • Updated • 406k • 1.59k • 1Note Synthetic speech for Indic languages
satwc-reddy/indian-language-deepfake-speech
Updated • 20Note Multilingual deepfake speech dataset for Indian languages (real+synthetic)