MiMo Audio Tokenizer (encoder only) -- GGUF
GGUF conversion of the encoder from XiaomiMiMo/MiMo-Audio-Tokenizer for use with CrispStrobe/CrispASR.
Available variants
| File | Quant | Size | Notes |
|---|---|---|---|
mimo-tokenizer-q4_k.gguf |
Q4_K | 377 MB | Encoder + RVQ codebooks |
Model details
- Architecture: 32-layer transformer encoder (1280d, 20 heads) + Conv1d stem + 20 RVQ codebooks
- Parameters: ~600M (encoder only, decoder/vocoder excluded)
- Audio: 24kHz input, outputs RVQ tokens at 25 Hz (8 channels used by ASR)
- License: MIT
- Source:
XiaomiMiMo/MiMo-Audio-Tokenizer
Notes
- Only the encoder is included (waveform โ RVQ tokens). Decoder/vocoder (TTS reconstruction) excluded.
- Used as the first stage of MiMo-V2.5-ASR pipeline (tokenizer โ LLM)
Provenance and EU AI Act Art. 53 note
- Upstream model: XiaomiMiMo/MiMo-Audio-Tokenizer โ published by
XiaomiMiMo. - Upstream licence:
mit. This repository redistributes under the same terms; it grants no rights the upstream licence does not. - What was done here: format conversion and/or quantisation only (GGUF/GGML). No training, no fine-tuning, no merging, no distillation, no change to architecture, vocabulary or capability. Only the numeric representation of the upstream weights differs.
- Training data: documented โ where it is documented at all โ by the upstream provider; see the upstream model card. No training data was used, added or selected by this repository.
- Provider status: under Regulation (EU) 2024/1689 the upstream authors remain the provider of this model. Converting the serialisation format does not make this repository the provider of a new general-purpose AI model, and no such claim is made. Questions about training content, copyright policy or model capability belong upstream.
- Downloads last month
- 707
Hardware compatibility
Log In to add your hardware
Model tree for cstr/mimo-tokenizer-GGUF
Base model
XiaomiMiMo/MiMo-Audio-Tokenizer