File size: 2,724 Bytes
3509f81
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
---
language: bn
license: unknown
tags:
- audio
- automatic-speech-recognition
- wav2vec2
- bengali
- bangla
- ctc
datasets:
- qdv206/wv-shru-v3-s6
---

# Bangla Wav2Vec2 (Unigram) — ASR Model

A Wav2Vec2-based CTC acoustic model for Bengali (Bangla) automatic speech recognition (ASR).

## Attribution

This model's weights and configuration were **not trained by the uploader**. They are mirrored from the Kaggle dataset
[`wv-shru-v3-s6`](https://www.kaggle.com/datasets/qdv206/wv-shru-v3-s6/data) published by Kaggle user
[**qdv206**](https://www.kaggle.com/qdv206). All credit for training and releasing the original model goes to the original author.
This upload exists to make the checkpoint easier to load and use via the `huggingface_hub` / `transformers` ecosystem.

If you use this model, please credit the original author and link back to the source dataset above.

## Model details

- **Architecture:** Wav2Vec2 (CTC head), `hidden_size=1024`, checkpoint tag `full_shru_v2_s20`
- **Task:** Automatic Speech Recognition (speech-to-text), Bengali script output
- **Vocabulary:** Character-level CTC vocabulary (90 tokens) covering Bengali script, digits, and basic punctuation
- **Sampling rate:** 16 kHz mono audio expected
- **Tokenizer/Processor:** `Wav2Vec2CTCTokenizer` + `Wav2Vec2FeatureExtractor` (`Wav2Vec2Processor`)

> Note: the original `config.json` lists `architectures: ["Wav2Vec2ForCTCV2"]`, a custom class name used in the
> original training pipeline. The underlying weights are a standard Wav2Vec2-for-CTC architecture, so the model
> loads with the standard `Wav2Vec2ForCTC` / `AutoModelForCTC` classes from `transformers` (see usage below).

## Usage

```python
import torch
import soundfile as sf
from transformers import Wav2Vec2ForCTC, Wav2Vec2Processor

model_id = "SayedShaun/bangla-wave2vec2-unigram"

processor = Wav2Vec2Processor.from_pretrained(model_id)
model = Wav2Vec2ForCTC.from_pretrained(model_id)

speech, sr = sf.read("audio.wav")  # expects 16kHz mono
inputs = processor(speech, sampling_rate=16000, return_tensors="pt", padding=True)

with torch.no_grad():
    logits = model(inputs.input_values).logits

predicted_ids = torch.argmax(logits, dim=-1)
transcription = processor.batch_decode(predicted_ids)
print(transcription)
```

## Source

- Original dataset/model: https://www.kaggle.com/datasets/qdv206/wv-shru-v3-s6/data
- Mirrored to Hugging Face for easier programmatic access.

## License

No explicit license was specified by the original author on Kaggle. Please refer to the
[original dataset page](https://www.kaggle.com/datasets/qdv206/wv-shru-v3-s6/data) for usage terms, and contact the
original author for clarification if you plan to use this commercially.