Buckets:
language:
- en
- zh
- fr
- it
- ko
- pt
- vi
license: mit
pipeline_tag: automatic-speech-recognition
tags:
- ASR
- quantization
- cpu-inference
- gguf
- bitnet
- multilingual
library_name: ggml
VibeVoice-ASR-BitNet
VibeVoice-ASR-BitNet is a compressed variant of VibeVoice-ASR optimized for real-time inference on edge CPUs — no GPU required. Through heterogeneous quantization, the model is compressed from 4.62 GB to 1.58 GB while achieving 1.6–2.3× faster inference than Whisper.cpp with real-time capability (RTF < 1) on as few as 3 CPU threads.
➡️ Code: microsoft/VibeASR.cpp
➡️ Report: VibeVoice-ASR-BitNet Technical Report
➡️ Base Model: microsoft/VibeVoice-ASR
🔥 Key Features
- ⚡ Real-time on CPU — RTF < 1 with 3+ threads on commodity x86 (AVX2) and ARM (NEON) hardware
- 📦 Compact — 1.58 GB total (2.9× compression from FP16), fits in edge device memory
- 🌍 Multilingual — English, Chinese, French, Italian, Korean, Portuguese, Vietnamese, and more
- 🔧 Custom SIMD Kernels — Fused operators within the ggml framework for both ARM and x86 platforms
Quantization Strategy
| Component | FP16 | Quantized | Method | Compression |
|---|---|---|---|---|
| VAE Tokenizer | 1.31 GB | 0.65 GB | I8_S | 2.0× |
| LM Decoder | 3.32 GB | 0.92 GB | I2_S + Q6_K | 3.6× |
| Total | 4.62 GB | 1.58 GB | — | 2.9× |
Evaluation
Inference Speed
| Threads | 1 | 2 | 3 | 4 | 6 | 8 |
|---|---|---|---|---|---|---|
| RTF | 1.98 | 1.08 | 0.77 | 0.63 | 0.49 | 0.42 |
| vs. Whisper.cpp | 2.28× | 2.12× | 1.86× | 1.86× | 1.71× | 1.55× |
Benchmarked on AMD EPYC 7V13 (AVX2+FMA) with 20s audio. Bold = RTF < 1 (real-time).
Accuracy (WER%)
| Benchmark | VibeVoice-ASR-7B | VibeVoice-ASR-BitNet | Parakeet | Whisper | SenseVoice | FunASR |
|---|---|---|---|---|---|---|
| MLC-EN | 7.82 | 8.25 | 8.40 | 13.57 | 12.39 | 11.36 |
| MLC-FR | 16.03 | 17.41 | — | — | — | — |
| MLC-IT | 15.67 | 17.23 | — | — | — | — |
| MLC-KO | 9.83 | 11.15 | — | — | — | — |
| MLC-PT | 22.41 | 24.87 | — | — | — | — |
| MLC-VI | 20.15 | 22.38 | — | — | — | — |
| AISHELL4 | 19.83 | 27.45 | — | — | 22.52 | 20.41 |
| AMI-ihm | 17.42 | 21.36 | 21.92 | 27.07 | 30.81 | 32.07 |
| AMI-sdm | 24.18 | 25.87 | 26.33 | 36.92 | 48.11 | 40.17 |
| AliMeeting | 36.21 | 40.58 | — | — | 38.75 | 39.27 |
| Fleurs-en | 4.73 | 5.21 | 4.09 | 3.99 | 6.84 | 4.93 |
| Fleurs-zh | 7.92 | 8.35 | — | — | 5.56 | 7.00 |
| Libri-clean | 2.17 | 2.41 | 1.49 | 1.98 | 2.78 | 1.58 |
| Libri-other | 5.84 | 6.27 | 3.13 | 3.60 | 6.81 | 4.01 |
| VoxPopuli | 4.92 | 5.18 | 5.26 | 7.19 | 8.63 | 6.46 |
Model Files
| File | Size | Description |
|---|---|---|
vibeasr-vae-encoder-i8_s.gguf |
0.65 GB | VAE tokenizer, I8_S quantized (ready to use) |
vibeasr-lm-i2_s-embed-q6_k.gguf |
0.92 GB | LM decoder, I2_S quantized (ready to use) |
model-*.safetensors |
10.7 GB | Original SafeTensors (for conversion) |
License
This project is licensed under the MIT License.
Contact
This project was conducted by members of Microsoft Research. If you have suggestions, questions, or observe unexpected behavior, please contact us at VibeVoice@microsoft.com.
Xet Storage Details
- Size:
- 4.12 kB
- Xet hash:
- aec2a419cf1c28cb0085c46d7acb180524a94a0e2c8543a171aa8787b4e2a67f
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.