| --- |
| license: apache-2.0 |
| base_model: mistralai/Voxtral-Mini-4B-Realtime-2602 |
| base_model_relation: quantized |
| pipeline_tag: automatic-speech-recognition |
| library_name: coreai |
| language: |
| - en |
| - fr |
| - es |
| - de |
| - ru |
| - zh |
| - ja |
| - it |
| - pt |
| - nl |
| - ar |
| - hi |
| - ko |
| tags: |
| - coreai |
| - aimodel |
| - apple-silicon |
| - voxtral_realtime |
| - speech |
| - audio |
| - streaming |
| --- |
| |
| # Voxtral Mini 4B Realtime β Core AI |
|
|
| Core AI conversions of [mistralai/Voxtral-Mini-4B-Realtime-2602](https://huggingface.co/mistralai/Voxtral-Mini-4B-Realtime-2602), made from source revision `2769294da9567371363522aac9bbcfdd19447add`. |
|
|
| These files require **Core AI 1.0.0b2** (macOS 27.0+, Xcode 27.0.0-beta.4+). They are **not** compatible with Core ML, MLX, GGUF/llama.cpp, or ONNX Runtime. **Use `voxtral_rt_decoder_int4_qkv_sdpa_cl512` + `voxtral_rt_encoder_stream_int4_buf3072.aimodel` + `voxtral_rt_encoder_first_chunk_float16_buf3072.aimodel` unless you have a specific size/speed requirement.** |
| |
| ## What this is |
| |
| Voxtral Mini 4B Realtime is Mistral's streaming speech-to-text model (Whisper encoder + MLP adapter + 26-layer Ministral-3 LLM decoder). This repository contains a complete native Core AI port: batch encoder (fp16), streaming encoder (int4), batched first-chunk encoder (fp16), and unified prefill+decode decoder (int4 with fused QKV projection and SDPA fused attention). |
| |
| The port achieves **4.45s end-to-end** for a 12s clip (RTF 0.37), **0.148s streaming TTFT** (time-to-first-token, compute-only), and **100% token parity** with the Hugging Face reference implementation in both offline and streaming modes. |
| |
| ## What will and will not load it |
| |
| | Runtime | Compatible? | |
| |---------|------------| |
| | Core AI 1.0.0b2+ (macOS 27.0+) | β
Yes | |
| | Core ML (`.mlmodelc`) | β No β different format | |
| | MLX | β No β different runtime | |
| | GGUF / llama.cpp / whisper.cpp | β No β different format | |
| | ONNX Runtime | β No β different format | |
| | transformers / PyTorch | β No β this is a compiled Core AI asset, not a checkpoint | |
| |
| ## Artifact manifest |
| |
| | File | Bytes | SHA-256 | Role | Recommended? | |
| |------|------:|---------|------|-------------| |
| | `artifacts/voxtral_rt_decoder_int4_qkv_sdpa_cl512/voxtral_rt_decoder_int4_qkv_sdpa_cl512.aimodel/main.mlirb` | 2,519,512,295 | `4ed29e2b363682c8c88b4087b1766735f32912b87815b61e0c35ece2d88c530d` | Decoder (int4 + fused QKV + SDPA) | **Yes** | |
| | `artifacts/voxtral_rt_decoder_int8_qkv_sdpa_cl512/voxtral_rt_decoder_int8_qkv_sdpa_cl512.aimodel/main.mlirb` | 4,032,607,471 | `5867ccfd9bb984b3e68d0e3550d28a3019dc0ffb75a332e034f182b2262b6fc1` | Decoder (int8 + fused QKV + SDPA) | Higher quality | |
| | `artifacts/voxtral_rt_decoder_fp16_qkv_sdpa_cl512/voxtral_rt_decoder_fp16_qkv_sdpa_cl512.aimodel/main.mlirb` | 6,869,584,654 | `61b88de4a5eaa16ea6bdeb289911283bf0e17c42611b05fa9b9fdebacd842e14` | Decoder (fp16 + fused QKV + SDPA) | Reference | |
| | `artifacts/voxtral_rt_encoder_stream_int4_buf3072.aimodel/main.mlirb` | 561,800,107 | `0777b124f14f5d909233e7c5c6d6b89eea2a84eb93790d02f3db4200144d49b6` | Streaming encoder (int4, 8-frame chunks) | **Yes** | |
| | `artifacts/voxtral_rt_encoder_first_chunk_float16_buf3072.aimodel/main.mlirb` | 1,991,831,559 | `e293cbccfd2bd6a401aff633478221df43d106d8f6a9f4884837a0090b746676` | Batched first-chunk encoder (fp16, 312 frames in one call) | **Yes** | |
| | `artifacts/voxtral_rt_encoder_float16.aimodel/main.mlirb` | 1,993,153,653 | `5323ea58dd4184ebc34717b468fdb34ae6070023c885b0079a6804ec136c314c` | Batch encoder (fp16, whole-clip) | Reference only | |
| | `voxtral-coreai-conversion.tar.gz` | 1,861,612 | `0142bb77847d59d68f0f1ad4807e614c5566de7026c70cc4de8fb3551e7f07b6` | Conversion source snapshot | For reproduction | |
|
|
| **Why two encoders?** The streaming encoder processes 8-mel-frame chunks incrementally (for real-time streaming). The batched first-chunk encoder processes the initial 312-frame chunk in one graph call (for fast TTFT). Both are needed: the batched encoder handles the first chunk, then the streaming encoder handles subsequent chunks. |
|
|
| ## Quantization comparison |
|
|
| All variants use weight-only quantization with `symmetric_with_clipping` per-block-32. The oracle is the upstream BF16 model's greedy token sequence on a 12.02s synthetic `say` audio fixture. |
|
|
| | Variant | Size | Batch e2e | Decode ms/tok | Tokens vs oracle | Recommendation | |
| |---------|-----:|----------:|--------------:|-----------------:|---------------| |
| | fp16 + fused QKV + SDPA | 6.4 GB | 7.93s | 45.3 | 161/161 | Reference | |
| | int8 + fused QKV + SDPA | 3.8 GB | 6.11s | 35.6 | 161/161 | Higher quality | |
| | **int4 + fused QKV + SDPA** | **2.3 GB** | **4.74s** | **28.1** | **161/161** | **Recommended** | |
| | int4 + fused QKV + SDPA + batched first chunk | 2.3 GB | **4.45s** | **26.6** | **161/161** | **Best (batch)** | |
|
|
| **Streaming (5-pass median, matched warmup):** |
|
|
| | Path | TTFT (compute) | chunk+decode | Tokens | |
| |------|---------------|-------------|--------| |
| | Serialized (39 calls) | 1.530s | 59.6ms | 128/128 | |
| | **Batched (1 call)** | **0.148s** | **59.6ms** | **128/128** | |
|
|
| The batched first chunk is **10.3x faster** for TTFT with zero chunk+decode cost. |
|
|
| **WER** (synthetic `say` audio references, not human labels): 0.000 (3s/6s), 0.071 (12s), 0.051 (20s). Token parity (161/161 batch, 128/128 streaming) proves port fidelity; WER measures transcription accuracy against the synthetic references. |
|
|
| ## Reproduce |
|
|
| Source model: `mistralai/Voxtral-Mini-4B-Realtime-2602@2769294da9567371363522aac9bbcfdd19447add` |
| Converter: `apple/coreai-models@af98e4d3bfcc7f797ff1d2c23251bf15f55dd72c` |
| Optimization: `apple/coreai-optimization@7ad2df3c2766636f8e76e3f9a2b53a5d9db4b33d` |
| Torch converter: `apple/coreai-torch@c89f6a44713249a12a84beec9f3e0cf2206ecc38` |
| Toolkit SHA-256: `0142bb77847d59d68f0f1ad4807e614c5566de7026c70cc4de8fb3551e7f07b6` |
|
|
| ```bash |
| # Set up a clean working directory |
| export WORKDIR=$(pwd)/voxtral-coreai-repro |
| mkdir -p $WORKDIR && cd $WORKDIR |
| |
| # Clone and pin the three dependency repos |
| git clone https://github.com/apple/coreai-models.git && cd coreai-models && git checkout af98e4d3bfcc7f797ff1d2c23251bf15f55dd72c && cd $WORKDIR |
| git clone https://github.com/apple/coreai-optimization.git && cd coreai-optimization && git checkout 7ad2df3c2766636f8e76e3f9a2b53a5d9db4b33d && cd $WORKDIR |
| git clone https://github.com/apple/coreai-torch.git && cd coreai-torch && git checkout c89f6a44713249a12a84beec9f3e0cf2206ecc38 && cd $WORKDIR |
| |
| # Download source model |
| hf download mistralai/Voxtral-Mini-4B-Realtime-2602 --local-dir $WORKDIR/Voxtral-Mini-4B-Realtime-2602 |
| |
| # Download and extract conversion toolkit |
| hf download harshav/Voxtral-Mini-4B-Realtime-CoreAI voxtral-coreai-conversion.tar.gz --local-dir $WORKDIR |
| cd $WORKDIR && tar xzf voxtral-coreai-conversion.tar.gz |
| |
| # Run oracle (produces golden tokens + embeddings) |
| uv run make_oracle.py --model $WORKDIR/Voxtral-Mini-4B-Realtime-2602 |
| |
| # Run streaming oracle (produces streaming golden tokens + chunk mels) |
| uv run make_oracle_stream.py --model $WORKDIR/Voxtral-Mini-4B-Realtime-2602 |
| |
| # Export encoder (batch) |
| PYTHONPATH=$WORKDIR/coreai-models/python/src:$WORKDIR/coreai-optimization/src uv run export_encoder.py --model $WORKDIR/Voxtral-Mini-4B-Realtime-2602 --dtype float16 |
| |
| # Export encoder (streaming) |
| PYTHONPATH=$WORKDIR/coreai-models/python/src:$WORKDIR/coreai-optimization/src uv run export_encoder_stream.py --model $WORKDIR/Voxtral-Mini-4B-Realtime-2602 --mode int4 --buf 3072 |
| |
| # Export encoder (batched first chunk) |
| PYTHONPATH=$WORKDIR/coreai-models/python/src:$WORKDIR/coreai-optimization/src uv run export_encoder_first_chunk.py --model $WORKDIR/Voxtral-Mini-4B-Realtime-2602 --dtype float16 --buf 3072 |
| |
| # Export decoder (int4 + fused QKV + SDPA) |
| PYTHONPATH=$WORKDIR/coreai-models/python/src:$WORKDIR/coreai-optimization/src uv run export_decoder.py --model $WORKDIR/Voxtral-Mini-4B-Realtime-2602 --mode int4 --fuse-qkv --use-sdpa |
| |
| # Gate (batch) |
| PYTHONPATH=$WORKDIR/coreai-models/python/src:$WORKDIR/coreai-optimization/src uv run gate_e2e.py \ |
| --decoder artifacts/voxtral_rt_decoder_int4_qkv_sdpa_cl512/voxtral_rt_decoder_int4_qkv_sdpa_cl512.aimodel \ |
| --compute-encoder gpu --compute-decoder ane |
| |
| # Gate (streaming, 5-pass median) |
| PYTHONPATH=$WORKDIR/coreai-models/python/src:$WORKDIR/coreai-optimization/src uv run gate_stream.py \ |
| --decoder artifacts/voxtral_rt_decoder_int4_qkv_sdpa_cl512/voxtral_rt_decoder_int4_qkv_sdpa_cl512.aimodel \ |
| --oracle golden_stream --buf 3072 --passes 5 \ |
| --first-chunk-encoder artifacts/voxtral_rt_encoder_first_chunk_float16_buf3072.aimodel |
| ``` |
|
|
| ## Usage |
|
|
| Download the recommended artifacts: |
|
|
| ```bash |
| hf download harshav/Voxtral-Mini-4B-Realtime-CoreAI \ |
| artifacts/voxtral_rt_decoder_int4_qkv_sdpa_cl512/voxtral_rt_decoder_int4_qkv_sdpa_cl512.aimodel/main.mlirb \ |
| artifacts/voxtral_rt_encoder_stream_int4_buf3072.aimodel/main.mlirb \ |
| artifacts/voxtral_rt_encoder_first_chunk_float16_buf3072.aimodel/main.mlirb \ |
| --local-dir ./Voxtral-Mini-4B-Realtime-CoreAI |
| ``` |
|
|
| Run with Core AI 1.0.0b2+ on macOS 27.0+ (Apple Silicon). The decoder uses a 512-token KV cache; the streaming encoder uses a 3072-frame buffer (~61s append-only capacity). Longer streams require rolling eviction (not implemented). |
|
|
| **Requirements:** |
| - macOS 27.0+ (Darwin 27.0.0+) |
| - Xcode 27.0.0-beta.4+ (for Core AI runtime) |
| - Apple Silicon (M4 Max validated; other Apple Silicon should work but unvalidated) |
| - ~4GB free disk space for the three recommended artifacts |
| - ~8GB unified memory during inference (model + KV cache + encoder buffers) |
|
|
| **Known limitations:** |
| - Streaming capacity is fixed at ~61s (buf=3072 append-only). Rolling KV eviction is not implemented. |
| - The batch encoder processes whole clips up to 12s. Longer clips require the streaming path. |
| - The decoder's `num_delay_tokens` is baked at 6 (480ms delay). Other delay values require re-export. |
| - ANE execution was attempted but the encoder's ANE compile failed (`ANECCompile: MLIR MPS to ANEC conversion failed`). The decoder with ANE preference runs ~10% faster but xctrace shows zero ANE hardware intervals β the preference changes codegen but doesn't execute on ANE. |
|
|
| ## Validation evidence |
|
|
| Machine-readable evidence is in `evidence/`: |
|
|
| - `conversion_gate.json` β golden token/embedding shapes from the oracle |
| - `benchmark.json` β batch and streaming benchmark results with protocol |
| - `quantization_quality.json` β WER against synthetic `say` references (not human labels) |
| - `artifact_hashes.json` β exact sizes and SHA-256 for all published artifacts |
| - `environment.json` β pinned revisions for model, converter, runtime, and toolchain |
|
|
| ## License and attribution |
|
|
| Upstream model: [mistralai/Voxtral-Mini-4B-Realtime-2602](https://huggingface.co/mistralai/Voxtral-Mini-4B-Realtime-2602) by Mistral AI, licensed [Apache-2.0](https://www.apache.org/licenses/LICENSE-2.0). |
|
|
| Base model: [mistralai/Ministral-3-3B-Base-2512](https://huggingface.co/mistralai/Ministral-3-3B-Base-2512) by Mistral AI. |
|
|
| Conversion: This repository is a Core AI conversion and does not claim authorship of the upstream model. The conversion toolkit (`voxtral-coreai-conversion.tar.gz`) is provided for reproduction. |
|
|
| Runtime: [Apple Core AI](https://github.com/apple/coreai-models) (BSD-3-Clause). |
|
|
| ## Rejected variants |
|
|
| The following were tested and rejected (not uploaded): |
|
|
| - **fp16 decoder** (6.4GB, 7.66s e2e) β too large, too slow. Reference only. |
| - **int8 decoder** (3.8GB, 6.24s e2e) β larger than int4, slower than int4+QKV+SDPA. No advantage. |
| - **int4 without QKV/SDPA** (2.3GB, 5.44s e2e) β slower than the fused variant. No advantage. |
| - **K/V state fusion** (52β26 writes) β engine segfault. Not viable. |
| - **Baked AdaRMSNorm constants** β 39.8 ms/tok (regression from 36.3). Index op dispatch worse than fused matmuls. |
| - **Broadcast GQA** β 34.1 ms/tok (marginal regression). Reshape overhead > repeat_interleave copy. |
| - **Fused write+fetch** (`mutable_cache_update_and_fetch`) β no improvement (34.6 vs 34.0 warm). |
| |