Papers
arxiv:2609.08672

X2Streaming-ASR: wait when uncertain, emit when ready for streaming ASR

Published on Oct 8
Authors:
,
,
,
,
,
,

Abstract

Streaming automatic speech recognition (ASR) for real-time voice agents and full-duplex dialogue must provide accurate partial transcripts with low commit latency. Existing systems commonly use a fixed chunk size, look-ahead, or target delay, or encourage emissions near estimated acoustic boundaries. These approaches do not directly optimize how much additional context to use at each output position under a single-pass, hard-commit constraint. We propose X2Streaming-ASR, which decomposes streaming recognition into when to commit and what to commit. Its three-stage training procedure first establishes streaming recognition ability, then warm-starts the commit policy with automatically probed trajectories, and finally refines the policy using character-level, segment-assigned group-relative rewards for recognition accuracy and latency. Across ten Chinese and English test sets, X2Streaming-ASR attains the lowest mean commit latency relative to forced-aligned endpoints, 32--109ms on Chinese characters and 12--85ms on English words, while recognition accuracy remains comparable to existing systems.

Community

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.08672
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 1

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.08672 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.08672 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.