File size: 2,526 Bytes
cf1fe95 3e4162a 8b6556c 3e4162a cf1fe95 3e4162a cac0e06 da768b4 cac0e06 3e4162a cac0e06 cf1fe95 3e4162a cf1fe95 3e4162a cf1fe95 8b6556c cf1fe95 3e4162a cf1fe95 3e4162a cf1fe95 8b6556c | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 | ---
language:
- ur
- sd
license: apache-2.0
pretty_name: Canopy-M
pipeline_tag: automatic-speech-recognition
library_name: pytorch
tags:
- automatic-speech-recognition
- speech
- ctc
- conformer
- urdu
- sindhi
---
# Canopy-M
<div align="center">
<img src="figures/canopy-m-logo.png" alt="Canopy-M Logo" width="220"/>
</div>
Canopy-M is a 57.9-million-parameter Conformer-CTC model for Urdu and Sindhi. It is non-autoregressive: a Macaron encoder predicts a token at every frame, and greedy CTC reads that sequence.

## Use
Install the library once, then point it at any published variant:
```bash
pip install torch torchaudio soundfile
pip install git+https://github.com/Proxima-AI-Co/canopy.git
hf auth login
canopy-transcribe clip.wav --model canopy-m --language urd
canopy-transcribe clip.wav --model ProximaAI/Canopy-M --language snd --device cpu
```
```python
from canopy import Canopy
asr = Canopy.from_pretrained("canopy-m")
print(asr.transcribe("clip.wav", language="urd"))
print(asr.transcribe("clip.wav", language="snd"))
```
## Training data
About **186.7 hours** were used in training.
| Language | Source | Train hours | Dev hours | Labels |
|---|---|---:|---:|---|
| Urdu | Common Voice + mahwiz + UrduSpeech | ~79.6 | ~17 | human |
| Sindhi | Common Voice | 71.5 | 0.10 | human; mostly unvalidated `other` |
| Sindhi | FLEURS `sd_in` | 12.3 | 1.33 | human, CC-BY-4.0 |
## Intended use
Research transcription of read Urdu and Sindhi at 16 kHz, one language per utterance, with the language given.
## Limitations
- Sindhi Common Voice development audio is 0.10 hours (70 utterances). Most of the Sindhi score is FLEURS.
- Common Voice Sindhi training audio is mostly unvalidated `other`, included with `--include-unvalidated`.
- Urdu Common Voice CER is 0.248. The headline 0.132 is pulled down by UrduSpeech and mahwiz.
- A wrong language id can cross Urdu and Sindhi, which share a Perso-Arabic script.
- No word n-gram was used in this table. Fusion can still overwrite a correct acoustic hypothesis.
## Citation
```bibtex
@software{canopy_m_2026,
title = {Canopy-M: Conformer-CTC speech recognition for Urdu and Sindhi},
author = {{Proxima AI}},
year = {2026},
url = {https://huggingface.co/ProximaAI/Canopy-M},
license = {Apache-2.0}
}
```
Also cite Common Voice, and FLEURS.
|