asr-324m-apache

A 7-language (Chinese, English, French, German, Japanese, Korean, Cantonese) speech-to-text model distilled from Qwen3-ASR-0.6B (Apache-2.0), depth- and vocabulary-pruned to 323.77M parameters β€” matching Audio8-ASR-0.1B's size, an existing model covering the same languages that is CC-BY-NC and cannot be used commercially. This checkpoint is fully Apache-2.0.

Base checkpoint (pre-compression, 467.81M, ties Audio8's quality): Luigi/asr-468m-apache-base. Code, full training pipeline, and every finding (including two negative results on adding more training data): github.com/vieenrose/asr-324m-apache.

Results (200-clip FLEURS test gate, all-refs; CER for zh/ja/ko/yue, WER for en/fr/de)

language Audio8-ASR-0.1B this model delta
French 20.88 19.29 βˆ’1.59 βœ…
Korean 13.97 12.47 βˆ’1.50 βœ…
Japanese 17.97 16.64 βˆ’1.33 βœ…
German 17.73 19.46 +1.73
English 8.51 12.36 +3.85
Chinese 11.94 18.00 +6.06
Cantonese 16.16 22.54 +6.38
macro 15.31 17.25 +1.94

3 of 7 languages individually beat Audio8 at this size; the overall average does not, yet β€” the remaining gap is a data-domain problem in Chinese and Cantonese specifically, not an architecture limit (see the linked repo's docs/findings.md).

Usage

This checkpoint's vocabulary is pruned (151,936 β†’ 42,000 ids), so decoding needs a remap step β€” vocab_remap.json (shipped in this repo) maps between the model's compact id space and the original Qwen3 tokenizer's ids.

import json
import torch
from huggingface_hub import hf_hub_download
from qwen_asr.core.transformers_backend.modeling_qwen3_asr import Qwen3ASRForConditionalGeneration
from qwen_asr.core.transformers_backend.processing_qwen3_asr import Qwen3ASRProcessor

path = "Luigi/asr-324m-apache"
proc = Qwen3ASRProcessor.from_pretrained(path)
model = Qwen3ASRForConditionalGeneration.from_pretrained(path, dtype=torch.bfloat16).cuda().eval()

remap = json.load(open(hf_hub_download(path, "vocab_remap.json")))
old_to_new = {int(k): v for k, v in remap["old_to_new"].items()}
keep_ids = remap["keep_ids"]
IM_END = 151645  # Qwen3 tokenizer's original <|im_end|> id

def transcribe(wav_16k_float32, language="Chinese", max_new_tokens=128):
    NATIVE = ("<|im_start|>system\n<|im_end|>\n<|im_start|>user\n<|audio_pad|><|im_end|>\n"
              "<|im_start|>assistant\n")
    e = proc(text=NATIVE + f"language {language}<asr_text>", audio=[wav_16k_float32],
             sampling_rate=16000, return_tensors="pt")
    e["input_ids"] = e["input_ids"].apply_(lambda i: old_to_new[i])
    eos = pad = old_to_new[IM_END]
    e = {k: (v.cuda() if torch.is_tensor(v) else v) for k, v in e.items()}
    if "input_features" in e:
        e["input_features"] = e["input_features"].to(torch.bfloat16)
    with torch.no_grad(), torch.autocast("cuda", dtype=torch.bfloat16):
        out = model.generate(**e, max_new_tokens=max_new_tokens, do_sample=False,
                             eos_token_id=eos, pad_token_id=pad)
    ids = out[0][e["input_ids"].shape[1]:].tolist()
    ids = [keep_ids[i] for i in ids]                    # map back to the original vocab space
    return proc.tokenizer.decode(ids, skip_special_tokens=True)

language accepts: Chinese, English, French, German, Japanese, Korean, Cantonese.

Training data and attribution

Trained on Common Voice 17 (CC0), WenetSpeech4TTS, Multilingual LibriSpeech, LibriSpeech, and FLEURS (all CC-BY-4.0). This model was trained in part on WenetSpeech4TTS, Multilingual LibriSpeech, LibriSpeech, and FLEURS, each licensed CC-BY-4.0 by their respective creators. Full source breakdown and every negative data-augmentation result are in the linked GitHub repository.

Audio8-ASR-0.1B is used only as a measurement reference throughout β€” its weights are never loaded, merged, or distilled from.

Downloads last month
-
Safetensors
Model size
0.3B params
Tensor type
F32
Β·
BF16
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for Luigi/asr-324m-apache

Finetuned
(49)
this model