Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level Timing
Abstract
Modern ASR models trained on heterogeneously annotated data treat transcription style (verbatim vs. intended) as an uncontrolled latent variable, causing measurable decoding instability, evaluation confounding (up to 60% of reported WER attributable to style mismatch), and unreliable word-level timing. We show that models already encode both styles; the challenge is controlled activation. Using coverage-aware decoder task tokens trained on parallel verbatim/intended transcript pairs, we raise German disfluency F1 from 10% to 79% zero-shot, despite English-only training. Full English-only fine-tuning surpasses all baselines in verbatim accuracy, disfluency detection, and intended-mode quality across both languages. We further introduce supervised cross-attention fine-tuning that improves word-level timestamps on disfluent speech beyond forced-alignment baselines. Finally, we propose verbatimize, a new task enabling scalable creation and enrichment of speech corpora with high-quality canonical verbatim transcriptions.
Community
Impactful paper for anyone working in transcription and TTS, Data and annotation pipelines. Data is notoriously unclean and most speech applications and evaluations suffer from it. This paper proposes a couple of interesting fixes and insights to solve this problem. We also introduced a couple of interesting improvements after publishing the paper, like mitigating hallucinations and a improved longform algorithm for Whisper style models. Read more here: https://nyra-labs.com/research
Get this paper in your agent:
hf papers read 2607.18934 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 4
nyralabs/CrisperWhisper2.0_large
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper