--- license: apache-2.0 language: - zh - yue - en pipeline_tag: automatic-speech-recognition library_name: transformers tags: - asr - speech-recognition - chinese - dialect - qwen3-asr - audio base_model: Qwen/Qwen3-ASR-1.7B base_model_relation: finetune ---

CN-MultiDialect-ASR logo

**On-Policy Self-Distillation for Multi-Dialect ASR: Mastering Dialects, Retaining Mandarin**

Shuiyuan Wang1 · Bingshen Mu1 · Pengshen Zhang2 · Chengyou Wang1 · Yujie Liao1 · Chengdong Liang2 · Binbin Zhang2 · Qiangze Feng3 · Lei Xie1

1 Audio, Speech and Language Processing Group (ASLP@NPU), School of Computer Science, Northwestern Polytechnical University, Xi'an, China
2 WeNet Community
3 NEXDATA TECHNOLOGY INC.

[![Paper](https://img.shields.io/badge/Paper-arXiv-blue)](https://arxiv.org/abs/2608.11898) [![GitHub](https://img.shields.io/badge/GitHub-ASLP--lab%2FCN--MultiDialect--ASR-black)](https://github.com/ASLP-lab/CN-MultiDialect-ASR) [![License](https://img.shields.io/badge/License-Apache%202.0-green)](https://huggingface.co/ASLP-lab/CN-MultiDialect-ASR)
This repository hosts the released **CN-MultiDialect-ASR** checkpoint, adapted from [Qwen3-ASR-1.7B](https://huggingface.co/Qwen/Qwen3-ASR-1.7B) with a three-stage pipeline: continual pre-training (CPT), dialect supervised fine-tuning (SFT), and On-Policy Self-Distillation (OPSD). The goal is to improve Chinese dialect recognition **without raising Mandarin CER**. - Paper: [arXiv:2608.11898](https://arxiv.org/abs/2608.11898) - Code, demo, and training scripts: [ASLP-lab/CN-MultiDialect-ASR](https://github.com/ASLP-lab/CN-MultiDialect-ASR)
OPSD framework

Overview of the staged adaptation pipeline. Top: base model, CPT, SFT, and OPSD. Bottom: OPSD with student on-policy prefixes, a frozen teacher conditioned on the reference transcript as privileged context, soft targets qt, and token-level KL.

## Demo Video demo with live waveforms and model transcriptions for Mandarin, English, four core dialects, and 15 ChinaVoices dialects. ## Key Features - **Mandarin–dialect balanced adaptation**: improves Chinese dialect ASR while retaining Mandarin recognition. - **Three-stage pipeline**: CPT strengthens the Chinese ASR foundation, dialect SFT specializes for dialects, and OPSD refines the final checkpoint. - **On-Policy Self-Distillation**: trains on student-decoded prefixes with soft teacher targets, reducing the train–test mismatch of teacher-forced ASR training. - **Drop-in inference**: compatible with the official [`qwen-asr`](https://github.com/QwenLM/Qwen3-ASR) package. ## Quickstart Inference is compatible with [Qwen3-ASR](https://github.com/QwenLM/Qwen3-ASR). We recommend installing the official `qwen-asr` package in a clean environment. ### Environment Setup ```bash conda create -n qwen3-asr python=3.12 -y conda activate qwen3-asr pip install -U qwen-asr ``` For faster inference with the vLLM backend: ```bash pip install -U qwen-asr[vllm] ``` ### Model Download You can load the model directly from Hugging Face, or download it locally first: ```bash # Hugging Face pip install -U "huggingface_hub[cli]" hf download ASLP-lab/CN-MultiDialect-ASR --local-dir ./CN-MultiDialect-ASR # ModelScope (recommended for users in Mainland China) pip install -U modelscope modelscope download --model ASLP-lab/CN-MultiDialect-ASR --local_dir ./CN-MultiDialect-ASR ``` ### Python Inference Load the model with `Qwen3ASRModel.from_pretrained` and call `transcribe`: ```python import torch from qwen_asr import Qwen3ASRModel model = Qwen3ASRModel.from_pretrained( "ASLP-lab/CN-MultiDialect-ASR", # or "./CN-MultiDialect-ASR" for a local path dtype=torch.bfloat16, device_map="cuda:0", # attn_implementation="flash_attention_2", max_inference_batch_size=32, max_new_tokens=256, ) results = model.transcribe( audio="path/to/audio.wav", language="Chinese", # or None for automatic language detection ) print(results[0].language) print(results[0].text) ``` Batch inference is also supported: ```python results = model.transcribe( audio=[ "path/to/mandarin.wav", "path/to/dialect.wav", ], language=["Chinese", "Chinese"], ) for r in results: print(r.language, r.text) ``` For vLLM backend, streaming inference, and forced alignment, see the [Qwen3-ASR repository](https://github.com/QwenLM/Qwen3-ASR). ## Method Overview | Stage | Training data | Goal | Objective | |:-----:|:--------------|:-----|:----------| | `CPT` | Full Mandarin-dialect collection (`~100k` hours) | Build a stronger Chinese ASR foundation | Cross-entropy | | `SFT` | Same sources with higher dialect sampling weight and a small Mandarin anchor | Lower dialect CER | Cross-entropy | | `OPSD` | Dialect refinement subset (`~5k` hours) | Improve dialect recognition without hurting Mandarin | Token-level KL | At inference time, only the student pathway is used. ## Performances ### Dialect Overview
Side-by-side radar of 1-CER on public and internal dialect sets

Higher is better. Left: 5 public dialect sets; right: 18 internal dialects. Both panels use the same radial scale (0.2–1.0). The figure compares the Qwen3-ASR baseline with the released CN-MultiDialect-ASR (OPSD) checkpoint.

### Public Dialect CER (%)
Evaluation set Dialect Qwen3-ASR CN-MultiDialect-ASR
WenetSpeech-Yue Long Cantonese 9.99 8.80
WenetSpeech-Yue Short Cantonese 6.93 5.31
WenetSpeech-Chuan Easy Sichuan 12.38 11.86
WenetSpeech-Chuan Hard Sichuan 21.79 21.74
WenetSpeech-Wu Wu 25.74 16.26
Dialect Avg. 15.37 12.79
### Internal Dialect CER (%)
Dialect Qwen3-ASR CN-MultiDialect-ASR
Anhui18.9513.08
Cantonese10.067.74
Changsha14.7910.23
Chaoshan45.5925.21
Dongbei6.455.80
Henan8.465.99
Kejia60.4728.60
Minnan30.0318.59
Nanchang33.4115.58
Nanjing13.379.33
Shanxi28.5318.69
Shaanxi9.686.28
Shandong8.787.64
Shanghai15.7812.07
Sichuan5.995.38
Suzhou50.3520.73
Wuhan11.307.59
Xuzhou6.125.04
Internal Avg.21.0112.42
### Mandarin CER (%)
Evaluation set Qwen3-ASR CN-MultiDialect-ASR
AISHELL-11.571.38
AISHELL-22.792.52
KeSpeech5.114.56
SpeechIO-10.750.86
SpeechIO-23.833.39
SpeechIO-31.391.27
Test_Meeting6.746.85
Test_Net5.465.30
Mandarin Avg.3.463.27
## Citation If you use this model, please cite: ```bibtex @misc{wang2026onpolicyselfdistillationmultidialectasr, title={On-Policy Self-Distillation for Multi-Dialect ASR: Mastering Dialects, Retaining Mandarin}, author={Shuiyuan Wang and Bingshen Mu and Pengshen Zhang and Chengyou Wang and Yujie Liao and Chengdong Liang and Binbin Zhang and Qiangze Feng and Lei Xie}, year={2026}, eprint={2608.11898}, archivePrefix={arXiv}, primaryClass={eess.AS}, url={https://arxiv.org/abs/2608.11898} } ``` ## License The released model is licensed under [Apache 2.0](https://www.apache.org/licenses/LICENSE-2.0). ## Contact For questions or collaborations, please contact [wangshuiyuan@mail.nwpu.edu.cn](mailto:wangshuiyuan@mail.nwpu.edu.cn). You are also welcome to join our WeChat group for technical discussions and updates.

WeChat group QR code
Scan to join our WeChat discussion group