Mizar-159M
Paper · Mizar family · GitHub code · Mizar-159M · Mizar-3B
Mizar is a 159.3M-parameter audio-language model for audio understanding. It combines a frozen CED-Small audio encoder, a frequency-merging audio–language mapper, and SmolLM2-135M. This repository releases the Stage-3 Q4 checkpoint, seed 20260905, update 200, trained with the joint strongAC + AVQA pool.
Code, environments, three-stage training and evaluation
Model at a glance
| Property | Value |
|---|---|
| Total parameters | 159,299,840 |
| Audio encoder | CED-Small, frozen |
| Mapper | Frequency-merging mapper, 126 audio tokens |
| Language decoder | SmolLM2-135M, fully fine-tuned |
| Input | One audio recording and an English question; optional answer choices |
| Audio preprocessing | Mono, 32 kHz, first 20 seconds, zero padding |
| Generation | Unrestricted greedy decoding, FP32, at most 300 new tokens |
| Checkpoint | Stage 3 Q4, seed 20260905, update 200 |
| Format | Full PyTorch tensor state dictionary; no optimizer state |
Benchmark accuracy
These are single-checkpoint results for the model in this repository.
| Benchmark | Accuracy (%) |
|---|---|
| MMAU full9k | 53.13 |
| MMAR (1,000) | 42.70 |
| ADQA-clean (1,577) | 36.02 |
| Unweighted three-benchmark mean | 43.95 |
MMAU full9k answers were generated locally and scored by the official service
(event ID 842faffae50c4d30bb0f10d418555f39).
MMAR and ADQA use their designated official evaluators. MMAU-mini is a
validation set and is not included in this table. Evaluation uses complete free
answers followed by official parsing; it does not rank option likelihoods or
constrain decoding to answer letters.
The inherited Stage-1 training data has known canonical audio matches to four MMAU full9k questions sharing one recording and 14 ADQA-clean questions. The fixed 1,577-question ADQA-clean set is the primary protocol; these results do not establish exhaustive absence of audio overlap or near duplicates.
Evaluation inputs and scoring code
- Prompts, the exact 9,000 model inputs:
data/evaluation/mmau_full9k_model_input.json - Answer parsing and scoring:
pipeline/format_predictions.py,pipeline/score_mmau.py,pipeline/score_adqa.py - Protocol:
docs/EVALUATION.md
Quick start
Clone the code and create the two environments:
git clone https://github.com/KaiyangLi1992/Mizar_159M.git
cd Mizar_159M
conda env create -f environment/train.yaml
conda env create -f environment/vllm.yaml
conda activate mizar-train
python pipeline/download_assets.py
python pipeline/verify_bundle.py
Create my_input.json:
[
{
"id": "my-audio",
"audio": "/absolute/path/to/recording.wav",
"question": "What can be heard in this recording?"
}
]
Then run audio-prefix preparation and vLLM answer generation:
python pipeline/prepare_inference.py \
--checkpoint checkpoints/final/stage3_q4_seed20260905.ckpt \
--input my_input.json --output inference/my_audio --device cpu
conda activate mizar-vllm
python pipeline/generate_vllm.py \
--bundle inference/my_audio \
--output inference/my_audio/answers.json
Prefix preparation may use CPU or CUDA; the supplied vLLM environment uses a CUDA GPU for decoding. The code preserves raw generated text, token IDs, finish reasons, audio hashes, checkpoint identity and engine version.
This is a composite audio-language architecture, not a standalone causal LM.
Load the checkpoint through the supplied code, not directly with
AutoModelForCausalLM.from_pretrained. The preparation command exports its
trained language decoder and continuous audio/text prefix for vLLM.
Training
- Stage 1: broad ReasonAQA + AudioMCQ audio instruction training, with 1,477,976 supervision rows in the historical executed mixture.
- Stage 2: audio-dependent fine-tuning on 256,077 strongAC and 33,875 AVQA rows; adopt update 906 from the original LR trajectory.
- Stage 3: Q4 refinement on the joint strongAC + AVQA pool, with cyclic answer-position balancing and teacher/student correctness obtained from free generation. Training uses ordinary gold-answer cross-entropy. The checkpoint is update 200 of the 600-update schedule.
Environment definitions, executable stage commands, exact data manifests, Q4 reconstruction, and protocol details are in the linked code repository. Training metadata can be downloaded with:
conda activate mizar-train
python pipeline/download_assets.py --training-data
python pipeline/unpack_data.py
Raw dataset audio and large S1/S2 CED feature caches are not included. Obtain source recordings from the dataset providers; the code documents path mapping and cache regeneration. S1/S2 intermediate weights are not part of this single-model release.
Files
checkpoints/final/stage3_q4_seed20260905.ckpt: complete Mizar weights.assets/: pinned CED-Small and SmolLM2 architecture/tokenizer/base assets.data/manifests/: compressed S1/S2/S3 candidate and training metadata.provenance/local_bundle.json: file hashes and model-weight identity.model_info.json: compact architecture and input/output metadata.
Intended use and limitations
Mizar is intended for research on compact audio understanding and audio question answering. It processes a single recording and does not use video. Only the first 20 seconds are represented; longer recordings require an application-level segmentation policy. The model can produce incorrect or unsupported descriptions and has primarily English text supervision.
License and attribution
The Mizar-159M weights and the files authored for this release are licensed
under the BSD 3-Clause Clear License. The model builds on CED-Small
and SmolLM2-135M, which are distributed under the Apache License 2.0; their
attribution and license text are kept in assets/UPSTREAM.md and
assets/LICENSE-APACHE-2.0. Dataset metadata under data/manifests/ is not
relicensed and remains subject to its upstream terms. The code repository
credits the Mellow implementation on which the training runtime is based.
See NOTICE.
Paper: Mizar: A 159M-Parameter Audio-Language Model for Audio Understanding, Kaiyang Li, Shaobo Han, Yue Tian, and Shihao Ji.
Model tree for KaiyangLi/Mizar-159M
Base model
HuggingFaceTB/SmolLM2-135M