Instructions to use addisai/addis-scribe-streaming with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- NeMo
How to use addisai/addis-scribe-streaming with NeMo:
import nemo.collections.asr as nemo_asr asr_model = nemo_asr.models.ASRModel.from_pretrained("addisai/addis-scribe-streaming") transcriptions = asr_model.transcribe(["file.wav"]) - Notebooks
- Google Colab
- Kaggle
Addis Scribe Streaming
Realtime Amharic speech recognition by Addis AI. Text appears while people speak, in 0.32-second steps, and is final about 30 ms after they stop. Built on NVIDIA Nemotron 3.5 ASR Streaming (0.6B), trained on 5,000+ hours of quality-checked Amharic speech.
| Addis Scribe Streaming | Google Chirp 3 (streaming) | |
|---|---|---|
| Word error rate, natural speech | 28.4% | 32.5% |
| Character error rate, natural speech | 13.1% | 16.5% |
| Final text after the speaker stops | 30 ms | 1,411 ms |
| First words on screen | 1.56 s | 6.67 s |
Streaming support
"Streaming" here means the model reads each new piece of audio once, keeps what it learned in a cache, and updates the text without re-reading earlier audio. Cost and delay stay flat however long someone talks.
| System | Streaming | How it handles live audio |
|---|---|---|
| Addis Scribe Streaming | Native | Cache-aware: each 0.32 s (or 1.12 s) chunk is processed once. |
| Google Chirp 3 | Native (cloud API) | Server-side streaming over the network. |
| Shook (Whisper medium) | Not supported | Whisper models transcribe complete audio, not a live stream. Our benchmark runs it live using the method of WhisperLiveKit, a widely used open-source tool for live Whisper: once per second it transcribes all audio received so far and displays a word once two consecutive transcriptions contain it. Because each pass covers the whole recording so far, the delay grows the longer someone speaks: measured median 10.8 s after speech ends. |
| Hohe (wav2vec2-BERT) | No | Offline only: needs the whole recording before it can answer. |
Benchmarks
Every system is scored with the same text normaliser (Unicode NFC, punctuation and symbols removed). WER = word error rate, CER = character error rate; lower is better.
Natural speech: 100 WAXAL Amharic test clips
22 speakers, 29 minutes, human transcripts. None of it was used in training. Clips were picked at random
(fixed seed) from WAXAL amh test, 2 to 30 s long.
| System | Mode | WER | CER | Final text after speech ends | First words |
|---|---|---|---|---|---|
| Addis Scribe Streaming | streaming 0.32 s | 28.4% | 13.1% | 30 ms | 1.56 s |
| Addis Scribe Streaming | streaming 1.12 s | 27.3% | 12.4% | 34 ms | 2.21 s |
| Google Chirp 3 (am-ET) | streaming, realtime | 32.5% | 16.5% | 1,411 ms | 6.67 s |
| Google Chirp 3 (am-ET) | batch, whole file | 29.2% | 13.5% | n/a | n/a |
| Shook (Whisper medium) | live, WhisperLiveKit method (1 s updates) | 35.0% | 16.6% | 10,839 ms | 3.66 s |
| Shook (Whisper medium) | offline | 35.0% | 16.6% | n/a | n/a |
| Hohe (wav2vec2-BERT) | offline only | 25.4% | 11.2% | n/a | n/a |
Latency for Addis Scribe and Shook is measured on one NVIDIA A100 with no network hop, with audio fed at real-time speed; Google's figures include the network round trip to its EU endpoint from the benchmark machine.
FLEURS
| System | Mode | WER | CER |
|---|---|---|---|
| Addis Scribe Streaming | streaming 0.32 s (first 150 clips) | 19.8% | 8.2% |
| Hohe | offline (same 150 clips) | 20.6% | 7.9% |
| Addis Scribe Streaming | offline (all 516 clips) | 20.6% | 7.7% |
| Hohe | offline (all 516 clips) | 20.2% | 7.3% |
| Previous Addis AI Nemotron | streaming 0.32 s (first 150 clips) | 35.0% | 17.0% |
Every per-clip transcript, latency and the scoring script are in benchmark/, so these
numbers can be reproduced exactly.
Quickstart
from stream_infer import AmharicStreamer
scribe = AmharicStreamer("addis-scribe-streaming.nemo", chunk="320ms") # or "1120ms"
for frame in microphone_frames(): # float32 mono, 16 kHz, any frame size
print(scribe.feed(frame)) # transcript so far
final_text = scribe.finish() # at end of speech; scribe.reset() for the next utterance
Command line: python stream_infer.py addis-scribe-streaming.nemo speech_16k.wav [320ms|1120ms].
Requires NVIDIA NeMo 3.0 with numba-cuda[cu12], numpy<2.4 and nvidia-nvjitlink-cu12>=12.9. Fed 100 ms frames
like a live microphone, AmharicStreamer reproduces the reference streaming transcripts exactly (100 of 100
benchmark clips).
Model
| Architecture | Cache-aware FastConformer encoder, hybrid RNNT + CTC, 0.6B parameters |
| Decoding | Greedy RNNT; language prompt am-ET (index 49) |
| Chunk sizes | 0.32 s (default) or 1.12 s |
| Input / output | 16 kHz mono audio / Amharic text without punctuation |
| Files | addis-scribe-streaming.nemo, serving.json, stream_infer.py, SHA256SUMS |
Training
Trained on 5,000+ hours of quality-checked Amharic speech, then refined on human-transcribed speech. FLEURS, WAXAL test and the evaluation speakers were never used in training.
Limits
- On natural conversational speech it is about 3 WER points behind Hohe, which is offline only.
- No punctuation (training text had punctuation removed).
- Tested on Amharic only; other languages are out of scope.
References
| Base model | NVIDIA Nemotron 3.5 ASR Streaming 0.6B |
| Toolkit | NVIDIA NeMo |
| Google Chirp 3 | Chirp 3 model documentation |
| Hohe | snapwre/hohe-asr-amharic |
| Shook | b1n1yam/shook-medium-amharic-2k |
| WhisperLiveKit | QuentinFuxa/WhisperLiveKit |
| WAXAL (natural-speech test set) | google/WaxalNLP, amh test split, CC-BY-SA-4.0 |
| FLEURS (read-speech test set) | google/fleurs, am_et test split |
License
Built on NVIDIA Nemotron 3.5 ASR Streaming 0.6B, used under the OpenMDW-1.1 license. © Addis AI.
- Downloads last month
- 160
8-bit
Model tree for addisai/addis-scribe-streaming
Base model
nvidia/nemotron-3.5-asr-streaming-0.6b