You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

faster-whisper-khasi

A fine-tuned Automatic Speech Recognition (ASR) model for the Khasi language, based on Whisper Large v3. This model is designed to transcribe spoken Khasi into text with improved accuracy on native speech compared to the original multilingual checkpoint.

Model Details

  • Model Name: faster-whisper-khasi
  • Base Model: openai/whisper-large-v3
  • Language: Khasi (kha)
  • Task: Automatic Speech Recognition (Speech-to-Text)

Model Description

faster-whisper-khasi is a fine-tuned version of Whisper Large v3 specifically adapted for Khasi speech recognition. The model has been trained to better understand Khasi pronunciation, vocabulary, and speaking styles while retaining Whisper's multilingual capabilities.

It is intended for applications such as:

  • Speech-to-text transcription
  • Voice assistants
  • Accessibility tools
  • Subtitle generation
  • Voice search
  • Educational applications
  • Language preservation projects

Training Data

The model was fine-tuned using the following dataset:

  • Dataset: toiar/Khasi_ASR_Dataset
  • Total Duration: Approximately 101 hours, 19 minutes, and 54.36 seconds
  • Language: Khasi
  • Task: Speech Recognition

Training

  • Base Model: openai/whisper-large-v3
  • Fine-tuning Task: Automatic Speech Recognition (ASR)
  • Input: Audio
  • Output: Khasi text transcription

Limitations

While the model performs significantly better on Khasi speech than the original multilingual model, it may still struggle with:

  • Heavy background noise
  • Strong regional accents not represented in the training data
  • Overlapping speakers
  • Code-switching between Khasi and other languages
  • Rare words or proper nouns
  • Low-quality audio recordings

Performance may vary depending on recording quality and speaking style.

Ethical Considerations

This model is intended to support the preservation and accessibility of the Khasi language. Users should ensure that speech data is collected and processed with appropriate consent and in accordance with applicable privacy regulations.

The model should not be relied upon where transcription errors could result in significant legal, medical, or safety consequences without human verification.

Citation

If you use this model in your research or application, please cite:

@misc{faster_whisper_khasi,
  title        = {faster-whisper-khasi},
  author       = {Toiar},
  year         = {2026},
  publisher    = {Hugging Face},
  howpublished = {\url{https://huggingface.co/toiar/faster-whisper-khasi}}
}
Downloads last month
6
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for toiar/faster-whisper-khasi

Finetuned
(916)
this model

Dataset used to train toiar/faster-whisper-khasi