Audio8-ASR-Infinite-MLX-4bit

MLX conversion of Edge0/Audio8-ASR-Infinite for Apple Silicon.

  • Converted from the base model at revision 7476824bc222e4ad509d286e8cae8b8d3f371129.
  • Not affiliated with or endorsed by Edge0.
  • Converted by Sasan Sotoodehfar, CAVI AI (https://cavi-ai.xyz).

Contents

  • model.safetensors: 3.74 GB (3.48 GiB).
  • Quantized to 4-bit (affine, group size 64): the text decoder (language_model.*), including the tied token embedding.
  • Kept in bf16: Voxtral Realtime audio tower, multi-modal projector, frame-length embedding, semantic VAD heads, ada_rms_norm MLPs.
  • audio8_asr_infinite/: MLX model code. mlx-audio 0.5.7 does not include this architecture; the snippet below registers it.

Requirements

  • Apple Silicon Mac.
  • Python 3.12.
  • pip install mlx-audio==0.5.7

Usage

import sys
from huggingface_hub import snapshot_download
path = snapshot_download("cavi-ai/Audio8-ASR-Infinite-MLX-4bit")
sys.path.insert(0, path)
import audio8_asr_infinite
sys.modules["mlx_audio.stt.models.audio8_asr_infinite"] = audio8_asr_infinite
from mlx_audio.stt.utils import load_model
model = load_model(path)
print(model.generate("audio.wav", language="en").text)
  • generate(audio, language, transcription_delay_ms=480): audio is a file path or a 16 kHz mono array.
  • language: "en" or "zh"; required.
  • transcription_delay_ms: positive multiple of 80; default 480.
  • Decoding: streaming greedy decode, one text token per 80 ms step.
  • Long audio: 30 s rolling decoder window.

Measured results

Hardware: Apple M5 Max.

Test 4-bit (this repo) bf16 weights, same MLX code
LibriSpeech validation-clean subset, hf-internal-testing/librispeech_asr_dummy (73 utterances, 1,150 words), WER 8.35% 7.30%
96 s English recording, WER 7.11% (decoded in 23.4 s)
Mandarin recording exact transcript
Peak memory, 6 s clip 4.2 GB
  • WER scoring: uppercase; punctuation removed except apostrophes; no number or spelling normalization.
  • Port check: MLX code vs a PyTorch fp32 reference built from transformers classes produced identical transcripts on 4 clips.

License

  • Model weights and configuration files: Apache-2.0, inherited from the base model. See LICENSE.
  • Code in audio8_asr_infinite/: MIT. See audio8_asr_infinite/LICENSE.

Links

Downloads last month
-
Safetensors
Model size
4B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for cavi-ai/Audio8-ASR-Infinite-MLX-4bit

Quantized
(3)
this model