Mizar-159M

Paper · Mizar family · GitHub code · Mizar-159M · Mizar-3B

Mizar is a 159.3M-parameter audio-language model for audio understanding. It combines a frozen CED-Small audio encoder, a frequency-merging audio–language mapper, and SmolLM2-135M. This repository releases the Stage-3 Q4 checkpoint, seed 20260905, update 200, trained with the joint strongAC + AVQA pool.

Code, environments, three-stage training and evaluation

Model at a glance

Property Value
Total parameters 159,299,840
Audio encoder CED-Small, frozen
Mapper Frequency-merging mapper, 126 audio tokens
Language decoder SmolLM2-135M, fully fine-tuned
Input One audio recording and an English question; optional answer choices
Audio preprocessing Mono, 32 kHz, first 20 seconds, zero padding
Generation Unrestricted greedy decoding, FP32, at most 300 new tokens
Checkpoint Stage 3 Q4, seed 20260905, update 200
Format Full PyTorch tensor state dictionary; no optimizer state

Benchmark accuracy

These are single-checkpoint results for the model in this repository.

Benchmark Accuracy (%)
MMAU full9k 53.13
MMAR (1,000) 42.70
ADQA-clean (1,577) 36.02
Unweighted three-benchmark mean 43.95

MMAU full9k answers were generated locally and scored by the official service (event ID 842faffae50c4d30bb0f10d418555f39). MMAR and ADQA use their designated official evaluators. MMAU-mini is a validation set and is not included in this table. Evaluation uses complete free answers followed by official parsing; it does not rank option likelihoods or constrain decoding to answer letters.

The inherited Stage-1 training data has known canonical audio matches to four MMAU full9k questions sharing one recording and 14 ADQA-clean questions. The fixed 1,577-question ADQA-clean set is the primary protocol; these results do not establish exhaustive absence of audio overlap or near duplicates.

Evaluation inputs and scoring code

Quick start

Clone the code and create the two environments:

git clone https://github.com/KaiyangLi1992/Mizar_159M.git
cd Mizar_159M
conda env create -f environment/train.yaml
conda env create -f environment/vllm.yaml
conda activate mizar-train
python pipeline/download_assets.py
python pipeline/verify_bundle.py

Create my_input.json:

[
  {
    "id": "my-audio",
    "audio": "/absolute/path/to/recording.wav",
    "question": "What can be heard in this recording?"
  }
]

Then run audio-prefix preparation and vLLM answer generation:

python pipeline/prepare_inference.py \
  --checkpoint checkpoints/final/stage3_q4_seed20260905.ckpt \
  --input my_input.json --output inference/my_audio --device cpu

conda activate mizar-vllm
python pipeline/generate_vllm.py \
  --bundle inference/my_audio \
  --output inference/my_audio/answers.json

Prefix preparation may use CPU or CUDA; the supplied vLLM environment uses a CUDA GPU for decoding. The code preserves raw generated text, token IDs, finish reasons, audio hashes, checkpoint identity and engine version.

This is a composite audio-language architecture, not a standalone causal LM. Load the checkpoint through the supplied code, not directly with AutoModelForCausalLM.from_pretrained. The preparation command exports its trained language decoder and continuous audio/text prefix for vLLM.

Training

  1. Stage 1: broad ReasonAQA + AudioMCQ audio instruction training, with 1,477,976 supervision rows in the historical executed mixture.
  2. Stage 2: audio-dependent fine-tuning on 256,077 strongAC and 33,875 AVQA rows; adopt update 906 from the original LR trajectory.
  3. Stage 3: Q4 refinement on the joint strongAC + AVQA pool, with cyclic answer-position balancing and teacher/student correctness obtained from free generation. Training uses ordinary gold-answer cross-entropy. The checkpoint is update 200 of the 600-update schedule.

Environment definitions, executable stage commands, exact data manifests, Q4 reconstruction, and protocol details are in the linked code repository. Training metadata can be downloaded with:

conda activate mizar-train
python pipeline/download_assets.py --training-data
python pipeline/unpack_data.py

Raw dataset audio and large S1/S2 CED feature caches are not included. Obtain source recordings from the dataset providers; the code documents path mapping and cache regeneration. S1/S2 intermediate weights are not part of this single-model release.

Files

  • checkpoints/final/stage3_q4_seed20260905.ckpt: complete Mizar weights.
  • assets/: pinned CED-Small and SmolLM2 architecture/tokenizer/base assets.
  • data/manifests/: compressed S1/S2/S3 candidate and training metadata.
  • provenance/local_bundle.json: file hashes and model-weight identity.
  • model_info.json: compact architecture and input/output metadata.

Intended use and limitations

Mizar is intended for research on compact audio understanding and audio question answering. It processes a single recording and does not use video. Only the first 20 seconds are represented; longer recordings require an application-level segmentation policy. The model can produce incorrect or unsupported descriptions and has primarily English text supervision.

License and attribution

The Mizar-159M weights and the files authored for this release are licensed under the BSD 3-Clause Clear License. The model builds on CED-Small and SmolLM2-135M, which are distributed under the Apache License 2.0; their attribution and license text are kept in assets/UPSTREAM.md and assets/LICENSE-APACHE-2.0. Dataset metadata under data/manifests/ is not relicensed and remains subject to its upstream terms. The code repository credits the Mellow implementation on which the training runtime is based. See NOTICE.

Paper: Mizar: A 159M-Parameter Audio-Language Model for Audio Understanding, Kaiyang Li, Shaobo Han, Yue Tian, and Shihao Ji.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for KaiyangLi/Mizar-159M

Finetuned
(948)
this model

Collection including KaiyangLi/Mizar-159M

Paper for KaiyangLi/Mizar-159M