Instructions to use x-square-robot/X2Streaming-ASR-4B-1009 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use x-square-robot/X2Streaming-ASR-4B-1009 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("automatic-speech-recognition", model="x-square-robot/X2Streaming-ASR-4B-1009")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("x-square-robot/X2Streaming-ASR-4B-1009") model = AutoModelForMultimodalLM.from_pretrained("x-square-robot/X2Streaming-ASR-4B-1009", device_map="auto") - Notebooks
- Google Colab
- Kaggle
X2Streaming-ASR-4B (zh/en)
Wait when uncertain, emit when ready: append-only streaming speech recognition for Chinese, English, and Chinese–English code-switching.
This checkpoint is X2Streaming-ASR-4B fine-tuned from Voxtral Mini Realtime with a learned commit policy (LISTEN / DECODE every 80 ms). Output is append-only—committed text is never rolled back.
| Architecture | VoxtralRealtimeForConditionalGeneration |
| Parameters | ~4B (26-layer text decoder, 32-layer audio encoder) |
| Precision | bfloat16 |
| Audio | 16 kHz mono; 128 mel bins; ~80 ms streaming frames (12.5 Hz) |
| Languages | zh, en, cs (code-switching) |
| Inference code | X-Square-Robot/X2Streaming-ASR |
Paper: X2Streaming-ASR (arXiv:2609.08672)
Model description
X2Streaming-ASR keeps the causal audio encoder, adapter, and decoder-only LM of Voxtral Realtime, but replaces fixed delay with a learned when to commit decision:
- LISTEN — For each new ~80 ms audio step, predict Wait if context is ambiguous, or Emit when ready.
- DECODE — After Emit, generate text for the new audio until EOS, then return to LISTEN.
- Reuse history — Incremental encoding and LM KV-cache reuse; no full-audio re-encoding each step.
On evaluated benchmarks (see the project README), mean emission latency is on the order of 32–109 ms (zh) and 12–85 ms (en) after each character/word ends, with competitive CER/WER versus streaming baselines.
Config summary (this repo)
| Component | Key settings |
|---|---|
| Text LM | 26 layers, hidden 3072, 32 heads, 8 KV heads, vocab 131072, sliding window 8192 |
| Audio encoder | 32 layers, hidden 1280, 32 heads, sliding window 750 |
| Streaming | downsample_factor: 4, audio_length_per_tok: 8, default_num_delay_tokens: 6 |
| Processor | VoxtralRealtimeProcessor, 16 kHz, hop 160, win 400 |
Files in this repository: config.json, params.json, processor_config.json, generation_config.json, tekken.json, and weight shards.
Quick start (recommended)
Weights alone are not enough for adaptive streaming—you need the X2Streaming-ASR inference stack (Transformers or vLLM).
1. Install
Linux, Python 3.11, and a CUDA GPU are recommended.
git clone https://github.com/X-Square-Robot/X2Streaming-ASR.git
cd X2Streaming-ASR
conda create -n x2streamingasr python=3.11 -y
conda activate x2streamingasr
pip install -r requirements.txt
2. Download this model
pip install -U huggingface_hub
# Optional mirror
export HF_ENDPOINT=https://hf-mirror.com
huggingface-cli download x-square-robot/X2Streaming-ASR-4B-1009 --local-dir ./X2Streaming-ASR-4B-1009
export MODEL_PATH="$(pwd)/X2Streaming-ASR-4B-1009"
Or clone from the Hub:
git lfs install
git clone https://huggingface.co/x-square-robot/X2Streaming-ASR-4B-1009
export MODEL_PATH=/path/to/X2Streaming-ASR-4B-1009
3. Streaming and offline inference
# Adaptive streaming (append-only commits)
python infer.py \
--config configs/infer.yaml \
--model_path "$MODEL_PATH" \
--language zh \
--audio /path/to/audio.wav
# Full-context offline
python infer.py \
--config configs/infer_offline.yaml \
--model_path "$MODEL_PATH" \
--language en \
--audio /path/to/audio.wav
Set
--languageto match your audio:zh(Chinese),en(English),cs(code-switching). One tag applies to the whole run.
4. JSONL batch
Input lines use audio, wav, or wav_path:
python infer.py \
--config configs/infer.yaml \
--model_path "$MODEL_PATH" \
--language en \
--input_file audio.jsonl \
--output_file results.jsonl
5. Web demo
BACKEND=transformers MODE=adaptive LANGUAGE=zh REPO_ID="$MODEL_PATH" bash demo/run.sh
Open http://localhost:7860.
Training note
This release supports Chinese and English recognition with the learned streaming commit policy described in the paper. Weights are initialized from Voxtral Mini Realtime and trained with the multi-stage recipe in X2Streaming-ASR (arXiv:2609.08672).
For full methodology and benchmarks, see the paper and the code repository.
License
Model weights: Apache License 2.0 (same as upstream Voxtral-Mini-4B-Realtime-2602). See the code repo for MODEL_LICENSE.md and attribution requirements.
Citation
@article{lin2026x2streaming,
title={X2Streaming-ASR: wait when uncertain, emit when ready for streaming ASR},
author={Lin, Zhiwei and Fu, Kaiqi and Wen, Rime and Liu, Zehan and Qin, Shawn and Gan, Roy and Wang, Hao and Wang, Qian},
journal={arXiv preprint arXiv:2609.08672},
year={2026}
}
简介(中文)
X2Streaming-ASR-4B 是面向中文、英文及中英混合的极低延迟、结果不可回退的流式语音识别模型,权重基于 Voxtral Mini Realtime 微调。
- 每约 80 ms 在 LISTEN(等待) 与 DECODE(发射) 间决策;已输出文本不会回改。
- 音频 16 kHz;请通过 X2Streaming-ASR 代码库 运行
infer.py或 Web Demo。 - 使用时务必设置 **
--language**:zh/en/cs。 - 论文:arXiv:2609.08672
- Downloads last month
- -
Model tree for x-square-robot/X2Streaming-ASR-4B-1009
Base model
mistralai/Ministral-3-3B-Base-2512