voidash's picture
correct Whisper NepTel figure 99.4 -> 96.3 (reproduces from published outputs)
dd7acc3 verified
|
Raw
History Blame Contribute Delete
2.69 kB
metadata
language:
  - ne
license: cc-by-nc-4.0
library_name: nemo
pipeline_tag: automatic-speech-recognition
tags:
  - automatic-speech-recognition
  - speech
  - nemo
  - conformer
  - streaming
  - telephony
  - nepali
  - nepal
model-index:
  - name: nepali-conformer-streaming
    results:
      - task:
          type: automatic-speech-recognition
        dataset:
          name: NepTel v0.1 (real Nepali call-center audio, human-reviewed)
          type: neptel
        metrics:
          - type: wer
            value: 59.87
            name: Real-call WER
          - type: cer
            value: 41.08
            name: Real-call CER
      - task:
          type: automatic-speech-recognition
        dataset:
          name: >-
            Held-out gold read Nepali (W1 read slice, OpenSLR-54 utterances
            absent from training)
          type: w1-read
        metrics:
          - type: wer
            value: 31.5
            name: Read-speech WER

nepali-conformer-streaming

Cache-aware streaming Nepali ASR (520 ms lookahead). Carries a large, honestly-reported streaming-lineage penalty on real calls — read RESULTS.md before choosing this over the offline model; it exists because a phone agent needs incremental output.

Try it: demo Space · Everything else: github.com/Ampixa/nepaliconformer (NepTel benchmark, per-system outputs, full honest results)

Numbers (measured, not marketed)

benchmark WER
NepTel — real Nepali call audio, human-reviewed refs 59.87
Held-out gold read Nepali (W1 slice) 31.5
Whisper-large-v3 zero-shot on the same NepTel audio 96.3

Architecture

121.3M-parameter 17-layer Conformer (d=512, striding ×4, 40 ms frames), hybrid TDT/CTC decoder, 1,024-piece Devanagari SentencePiece. Chunked-limited attention [[70,13],[70,6],[70,1],[70,0]], fully causal convolutions, cache-aware incremental decoding.

Training data

~1,655 h of mostly conversational Nepali (YouTube podcasts/interviews) with Google Chirp 2 pseudo-labels + 105 h human-labeled read speech; telephony codec, noise, reverb and tempo augmentation. Label-noise ceiling and every measured limitation (English, sung speech, slow speech, end-of-turn) are documented in the repo's RESULTS.md.

Usage

from nemo.collections.asr.models import EncDecHybridRNNTCTCBPEModel
m = EncDecHybridRNNTCTCBPEModel.restore_from("nepali_conformer_streaming.nemo")
print(m.transcribe(["audio.wav"])[0].text)

License: CC-BY-NC-4.0 (weights). Code in the repo: MIT.