File size: 2,526 Bytes
cf1fe95
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3e4162a
 
 
 
8b6556c
3e4162a
cf1fe95
 
 
 
 
3e4162a
cac0e06
 
da768b4
 
cac0e06
 
3e4162a
cac0e06
cf1fe95
 
3e4162a
cf1fe95
3e4162a
cf1fe95
 
 
 
 
 
 
8b6556c
cf1fe95
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3e4162a
cf1fe95
3e4162a
cf1fe95
 
 
 
8b6556c
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
---

language:
  - ur
  - sd
license: apache-2.0
pretty_name: Canopy-M
pipeline_tag: automatic-speech-recognition
library_name: pytorch
tags:
  - automatic-speech-recognition
  - speech
  - ctc
  - conformer
  - urdu
  - sindhi
---


# Canopy-M

<div align="center">
  <img src="figures/canopy-m-logo.png" alt="Canopy-M Logo" width="220"/>
</div>

Canopy-M is a 57.9-million-parameter Conformer-CTC model for Urdu and Sindhi. It is non-autoregressive: a Macaron encoder predicts a token at every frame, and greedy CTC reads that sequence.


![Canopy-M versus Whisper on the Urdu and Sindhi development sets](figures/overview_cer_wer.png)

## Use

Install the library once, then point it at any published variant:

```bash

pip install torch torchaudio soundfile

pip install git+https://github.com/Proxima-AI-Co/canopy.git

hf auth login

canopy-transcribe clip.wav --model canopy-m --language urd

canopy-transcribe clip.wav --model ProximaAI/Canopy-M --language snd --device cpu

```

```python

from canopy import Canopy



asr = Canopy.from_pretrained("canopy-m")

print(asr.transcribe("clip.wav", language="urd"))

print(asr.transcribe("clip.wav", language="snd"))

```


## Training data

About **186.7 hours** were used in training.

| Language | Source | Train hours | Dev hours | Labels |
|---|---|---:|---:|---|
| Urdu | Common Voice + mahwiz + UrduSpeech | ~79.6 | ~17 | human |
| Sindhi | Common Voice | 71.5 | 0.10 | human; mostly unvalidated `other` |
| Sindhi | FLEURS `sd_in` | 12.3 | 1.33 | human, CC-BY-4.0 |


## Intended use

Research transcription of read Urdu and Sindhi at 16 kHz, one language per utterance, with the language given.

## Limitations

- Sindhi Common Voice development audio is 0.10 hours (70 utterances). Most of the Sindhi score is FLEURS.
- Common Voice Sindhi training audio is mostly unvalidated `other`, included with `--include-unvalidated`.
- Urdu Common Voice CER is 0.248. The headline 0.132 is pulled down by UrduSpeech and mahwiz.
- A wrong language id can cross Urdu and Sindhi, which share a Perso-Arabic script.
- No word n-gram was used in this table. Fusion can still overwrite a correct acoustic hypothesis.

## Citation

```bibtex

@software{canopy_m_2026,

  title = {Canopy-M: Conformer-CTC speech recognition for Urdu and Sindhi},

  author = {{Proxima AI}},

  year = {2026},

  url = {https://huggingface.co/ProximaAI/Canopy-M},

  license = {Apache-2.0}

}

```

Also cite Common Voice, and FLEURS.