Instructions to use ARTPARK-IISc/Vaani-FastConformer-Malayalam with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- NeMo
How to use ARTPARK-IISc/Vaani-FastConformer-Malayalam with NeMo:
# tag did not correspond to a valid NeMo domain.
- Notebooks
- Google Colab
- Kaggle
Important...!!!!!
Thanks for this particular model ♥️
But it still leaves some lacuna, mispronounced words are there still..
Also the model is very high in size which needs lots of computing and skill to make it run for common people..
Still after all these developments, malayalam still lacks good ASR model.
All are proprietary handled by big firms, open ones are still far behind the game. People are still struggling. Malayalam in all sector is far behind(In TTS , LLM..etc) Just beacuse there are no Good Open ASR.
If one is decided to make a good model, still computing hits hard, training costs are high. Please make small one, one which is able to run via browser via ONNX, i suggest a below 800mb model. If its below 500mb, all welcome. I have seen proprietary firms having SOTA models below 450mb. For complex scripts like malayalam, it seems fastConformer TDT works well, try fastConformer-TDT-CTC hybrid for SOTA. Or fastConformerRNNT some decent production level model. CTC alone not working well.
If you are really into helping the people for everyday tasks, consider this suggestion, you'll be remembered...
Release in open type licence
- a unfancy researcher guy
Thanks for the feedback, the current model is 430 million parameter model with hybrid CTC -TDT decoder. Could you please share performance numbers(WER,CER) on your evaluation data for both opensource models as well proprietary models? we would like to see the performance gap.
thanks for the reply, my feedback is not from any benchmarking or proper evaluation, i was just recently (at least 2 years) started looking into ASR (especially ml) as it useful for my theses writing. I tested "ARTPARK-IISc/Vaani-FastConformer-Malayalam" for some random long audios, it struggles on a small scale at some point, where seperate words are mixed together as one word. and some words are spelled wrong (that too on minor scale). but its there. if audio is long enough like 20 minutes, it really need to scrutinisation to correct it, which is time consuming. i dont know how benchmark this, if i had known, i would have done it. Just consider my feedback as raw end users one, where manually verifying models output. one of the best model i have seen across my "ml model search" is ADALAT AI's https://huggingface.co/spaces/adalat-ai/malayalam-dictation. they're saying its a "Hybrid-TDT-CTC-512-L-Malayalam-v1". hoping its a FastConformer TDT-CTC. They said they are not planning to release that model. I think the model is below 500mb (like ai4bharath's IndicConformer) which is a game changer. it would be helpfull a smaller version is there , a below 1GBModel ( i would say below 500MB). If you really intent to help normal peoples please try to release a smaller model which is not much resource hungry..
I am still recommending single language specific models than multi lang models for individual use cases like mine (research in malayalam language ). There is still a lacuna in this sector. Real struggle is we people still rely on paid/proprietary solutions for our academic transcriptions. There are no privacy focused solutions. As of now there are many ML models out there, but still seems lagging behind. Am using mac, plethora of main stream models doesn't run in mac natively (like nemo). It needs MLX (.mlmodel) support for apple's MPS. Mac seems to have plenty vibe coded transcription apps out there, none of them have a realible indic language model except Hindi. But other indic language models seems appearing in past 1 year on HF. Still no malayalam yet. We need really a SOTA.
And about whisper, there are fasterwhisper, indicwhisper, vegam whisper....10000 whisper malayalam models are there,
All of them still are broken.
Also it goes same for TTS malayalam (0 good models, even after many small param game changers like kokoro, kyutai,suprrtonic,pockeTTS,piper,, soprano...) All we have is 2 ai4bharath heavy models (parler & F5 which still struggles for malayalam (robotic, 60% good compared to current standards). Just mentioned as if, it is useful for Artpark, if investing in TTS models in future.