--- license: apache-2.0 base_model: Qwen/Qwen3-ForcedAligner-0.6B pipeline_tag: automatic-speech-recognition library_name: openasr tags: - openasr - oasr - qwen3-forced-aligner-0.6b ---
# Qwen3-ForcedAligner 0.6B ยท OpenASR **Word-level forced alignment for OpenASR transcripts -- a non-autoregressive Qwen3 audio+text model that refines per-word timestamps** [![License](https://img.shields.io/badge/license-Apache--2.0-2563eb.svg)](https://huggingface.co/Qwen/Qwen3-ForcedAligner-0.6B/blob/c7cbfc2048c462b0d63a45797104fc9db3ad62b7/LICENSE) [![Format](https://img.shields.io/badge/format-.oasr-7c3aed.svg)](https://github.com/QuintinShaw/openasr) [![Runtime](https://img.shields.io/badge/runtime-OpenASR-111827.svg)](https://openasr.org) [![Base model](https://img.shields.io/badge/base-Qwen3--ForcedAligner--0.6B-f59e0b.svg)](https://huggingface.co/Qwen/Qwen3-ForcedAligner-0.6B) A capability-pack support model for the **[OpenASR](https://github.com/QuintinShaw/openasr)** runtime โ€” pure-Rust inference, **no Python at inference time**. Not a standalone transcription model: it augments another OpenASR ASR model's own decode path.
--- ## โœจ Highlights - ๐ŸŽฏ **Refined word timestamps** โ€” consumes a finished transcript's text plus the source audio and replaces a model family's own approximate per-word timestamps with aligner-refined spans (`--word-timestamps=aligned`) - โšก **Non-autoregressive** โ€” a single forward pass over interleaved audio/text with argmax at `` positions (5000 80ms-wide bins), not incremental greedy decoding, so it is not dispatched through the qwen3-asr runtime - ๐Ÿงฉ **Shares its backbone with Qwen3-ASR** โ€” the same audio-encoder + LM `thinker` tensor layout, byte-for-byte; only the final head differs (an independent 5000-way classification head instead of the tied vocabulary head) - ๐Ÿ”Œ **Shared attribution dependency** โ€” used explicitly by `--word-timestamps=aligned` and internally when external diarization must split a coarse ASR segment at speaker changes - ๐Ÿฆ€ **Validated native Q4_K default** โ€” the public q4_k name follows OpenASR's unified tier naming, while boundary-sensitive audio, token-embedding, and timestamp-head matrices stay Q8_0; Q8_0 and FP16 remain available - ๐Ÿฆ€ **Native in OpenASR** โ€” `.oasr` packs run with no Python at inference, engineered for peak performance on CPU & GPU ## ๐Ÿš€ Quickstart ```bash # 1. Install the OpenASR CLI ยท https://openasr.org # 2. Pull the pack openasr pull qwen3-forced-aligner-0.6b:q4 # 3. Use it as an opt-in refinement for another model's transcribe call openasr transcribe meeting.wav --model --word-timestamps=aligned ``` ## ๐Ÿ“ฆ Pack | Quant | File (`.oasr`) | Size | |:------|:---------------|-----:| | fp16 | `qwen3-forced-aligner-0.6b-fp16.oasr` | 1.84 GB | | q8_0 | `qwen3-forced-aligner-0.6b-q8_0.oasr` | 986 MB | | q4_k | `qwen3-forced-aligner-0.6b-q4_k.oasr` | 765 MB | ## ๐Ÿง  About Qwen3-ForcedAligner 0.6B Qwen3-ForcedAligner-0.6B is a **word-level forced-alignment** model from **Qwen**, sharing its audio-encoder + LM `thinker` tensor layout byte-for-byte with **Qwen3-ASR** (same `Qwen3ASRForConditionalGeneration` architecture). The only structural difference is the final head: instead of a tied vocabulary `lm_head`, it uses an independent `Linear(hidden_size, 5000)` classification head over 80ms-wide timestamp bins. Given a transcript's text and its source audio, it runs a single non-autoregressive forward pass and reads off word-boundary timestamps at argmax `` positions -- refining a model family's own (typically decode-time-approximate) per-word timestamps. This OpenASR repo repackages the weights as `.oasr` packs that run natively in the OpenASR runtime -- no Python at inference, all decoding local. OpenASR recommends the validated **q4_k** tier and also ships **q8_0** and an **fp16** full-precision reference tier. The q4_k label is the catalog's unified product name, not a claim that every matrix uses Q4_K: the audio encoder, token embedding, and timestamp head remain Q8_0, and pack verification replays the exact per-tensor policy. Q3 and legacy all-Q4 packs are rejected because small logit perturbations can move a word boundary across multiple 80ms bins. **Not a standalone transcription model.** This pack cannot transcribe audio by itself; it is an alignment dependency consumed explicitly via `openasr transcribe