Instructions to use leope/ark-asr-3B-mlx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use leope/ark-asr-3B-mlx with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir ark-asr-3B-mlx leope/ark-asr-3B-mlx
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
Validation
Validation was run on 2026-08-01 with an Apple M5 Pro, 24 GB unified memory, macOS 26.6, Python 3.12.13, and MLX 0.32.0.
Checkpoint conversion
- Source:
AutoArk-AI/ARK-ASR-3B - Revision:
1e28271b79edc97635783bea65abc89195a09ed3 - Source tensors: 926
- Intentionally dropped tensors: 2
- Retained parameter tensors: 924
- Missing, unexpected, or shape-mismatched tensors: 0
- Output dtype: BF16
- Output size: 7.0 GiB
- Output SHA-256:
a5e9431bdd648340a40c092e385c4fe1d445d8ad3311dd94609e72af36b0256d
Source shard hashes are recorded in conversion.json.
PyTorch parity
The original remote-code model and native MLX model were run on the pinned
Narsil/asr_dummy/1.flac LibriSpeech sample.
- Prompt token IDs: identical
- BF16 input features: maximum absolute difference
0.0 - Adapted audio feature cosine similarity:
0.998567558665821 - Initial decoder logits cosine similarity:
0.9999223476845636 - Greedy generation token IDs: identical
- Final decoded text: identical
Both implementations produced:
he hoped there would be stew for dinner turnips and carrots and bruised potatoes and fat mutton pieces to be ladled out in thick peppered flour fattened sauce
The validation command was:
python scripts/validate_parity.py /path/to/1.flac \
--model . \
--source AutoArk-AI/ARK-ASR-3B \
--revision 1e28271b79edc97635783bea65abc89195a09ed3 \
--min-adapter-cosine 0.998
MLX benchmark
The same 12.1-second audio sample was measured after the checkpoint was already present in the operating-system file cache:
- Model load:
1.011 s - Audio preprocessing:
0.499 s - First token:
0.885 s - Full 34-token generation:
1.965 s - Generation throughput:
17.30 tokens/s - MLX peak memory:
7.820 GB
These numbers describe this machine and test clip; they are not portable performance guarantees.