Canopy-M
Canopy-M is a 57.9-million-parameter Conformer-CTC model for Urdu and Sindhi. It is non-autoregressive: a Macaron encoder predicts a token at every frame, and greedy CTC reads that sequence.
Use
Install the library once, then point it at any published variant:
pip install torch torchaudio soundfile
pip install git+https://github.com/Proxima-AI-Co/canopy.git
hf auth login
canopy-transcribe clip.wav --model canopy-m --language urd
canopy-transcribe clip.wav --model ProximaAI/Canopy-M --language snd --device cpu
from canopy import Canopy
asr = Canopy.from_pretrained("canopy-m")
print(asr.transcribe("clip.wav", language="urd"))
print(asr.transcribe("clip.wav", language="snd"))
Training data
About 186.7 hours were used in training.
| Language | Source | Train hours | Dev hours | Labels |
|---|---|---|---|---|
| Urdu | Common Voice + mahwiz + UrduSpeech | ~79.6 | ~17 | human |
| Sindhi | Common Voice | 71.5 | 0.10 | human; mostly unvalidated other |
| Sindhi | FLEURS sd_in |
12.3 | 1.33 | human, CC-BY-4.0 |
Intended use
Research transcription of read Urdu and Sindhi at 16 kHz, one language per utterance, with the language given.
Limitations
- Sindhi Common Voice development audio is 0.10 hours (70 utterances). Most of the Sindhi score is FLEURS.
- Common Voice Sindhi training audio is mostly unvalidated
other, included with--include-unvalidated. - Urdu Common Voice CER is 0.248. The headline 0.132 is pulled down by UrduSpeech and mahwiz.
- A wrong language id can cross Urdu and Sindhi, which share a Perso-Arabic script.
- No word n-gram was used in this table. Fusion can still overwrite a correct acoustic hypothesis.
Citation
@software{canopy_m_2026,
title = {Canopy-M: Conformer-CTC speech recognition for Urdu and Sindhi},
author = {{Proxima AI}},
year = {2026},
url = {https://huggingface.co/ProximaAI/Canopy-M},
license = {Apache-2.0}
}
Also cite Common Voice, and FLEURS.
