Bonsai 27b MLX Fast

Benchmark at a glance

M4 Mac mini · 16 GB memory · short prompt · model already loaded.

Version Output tokens/sec Speed vs original First token Task checks
Original 1-bit Bonsai, stock-style settings 19.7 1.00× 1.62 s Not separately checked
Optimization 19.8 1.00×; no meaningful gain 1.60 s 15/18
Fast 22.0 1.12× 1.36 s 11/18
UltraFastUnStable, revised 24.6 1.25× 1.35 s 11/18, earlier same-candidate check

Read this before comparing: these are three-run medians for the same 81-token input and 128 output tokens, with thinking off and a tiny warm-up. Loading time is excluded. The revised UltraFastUnStable was measured later with a different memory ceiling, so its ratio is a comparison of recorded results—not a fresh same-session, controlled speedup. Longer inputs and tool schemas take longer. A higher rate does not mean better or correct answers; task counts are not percentages of intelligence retained.

The retired 48-block experiment reached 79.5 tokens/sec / 0.427 s TTFT, but produced nonsense or blank lines. It is unusable, is not the current UltraFastUnStable download, and is kept only in its historical revision tag. Measurement details and limitations.

Quick tour: modestly faster, less reliable

  1. What it is: ten original blocks physically removed; 54 remain. Its remaining weights are unchanged. It saves about 600 MB and some generation time, but lost four task passes versus Optimization.
  2. What to download: the entire repository through Files and versions or the CLI command below. Pruning is inside the checkpoint, but ritual_model.py and Prism's runtime are still required—no speed flag needed.
  3. How to try it: follow Get started with a short, low-stakes request. Check facts and validate tool calls. For strict JSON or dependable instruction following, choose Optimization instead.

A smaller, slightly faster—but less reliable—edition of Bonsai for Mac. This edition removes ten of the original model's 64 layers. Its remaining weights are unchanged. Removing layers saves work and storage, but it also changes the model's answers.

This is an experimental speed/quality tradeoff, not a better model overall. For reliable tool calls, factual answers, strict JSON, and following instructions, choose Bonsai 27b MLX Optimization.

The difference in plain English

Optimization Fast — this page
Model layers All 64 kept Ten removed; 54 kept
Weight-file size About 5.13 GB About 4.53 GB
Measured output speed 19.8 tokens/second 22.0 tokens/second
Time to first output token 1.60 seconds 1.36 seconds
Local task check 15/18 passed 11/18 passed

The speed benefit was about 2.3 extra tokens per second and 0.24 second less waiting for the first token. That is a modest improvement, not a huge speed jump. The weight file is about 600 MB smaller, and peak MLX memory was 4.06 GB instead of 4.71 GB.

Read the quality warning before downloading

In our small 18-task check, Fast passed four fewer tasks than Optimization: a 22 percentage-point drop in task pass rate. That is not a measurement of its overall intelligence, but it is a clear warning for the workloads tested.

Failures included a wrong multiplication result, invalid JSON, inventing an 80% success rate that was not present in the supplied information, and refusing to schedule rather than asking for missing details. A recursion explanation also became repetitive. Both editions failed some clarification requests.

Fast has not demonstrated the project's intended low-loss quality tradeoff. The earlier 13% tolerance was not an overall intelligence measurement, and this small task check cannot establish such a percentage. Fast should not be used when dependable facts, valid tool payloads, or strict instructions are important. Later experimental variants reached around 24–26 tokens/second but also failed capability checks, so they were rejected and are not included in this download.

See the detailed test results for the task checklist, timing conditions, and rejected experiments.

What the speed numbers mean

Tests used an M4 Mac mini with 16 GB memory, the same 81-token input, and 128 generated tokens per run. The table uses medians of three runs, with thinking off and greedy decoding. These are actual output-generation rates, not hidden-reasoning counts.

The model was already loaded and warmed up. Loading is not included in time to first token, and longer inputs can take much longer. These numbers are not a promise of a two-second response for every request or every Mac. Peak MLX memory is not the same as total system memory.

Important: one weight file is not enough

Use Files and versions and download the entire repository: weights, configuration, tokenizer, ritual_model.py, installer, and launch code. A single model.safetensors file does not contain the complete optimized setup. An app's “Download MLX” button does not install this required code or runtime for you.

The layer removal is already part of this checkpoint; no extra speed-mode flag is needed. However, the included custom loader is required to use these weights correctly, and Prism ML's MLX fork supplies the 1-bit kernels. This is not a universal drag-and-drop checkpoint for every MLX app. Use only model code you trust.

Get started

You need an Apple Silicon Mac, uv, and Apple's Metal toolchain. The installer builds the required runtime with one compilation job; compilation still takes time and space. The compiled runtime is not included.

hf download VertexAIco/Bonsai-27b-MLX-Fast --local-dir Bonsai-27b-MLX-Fast
cd Bonsai-27b-MLX-Fast
./install_prism_mlx.command
.venv-1bit/bin/python run_bonsai_1bit.py --interactive

Sign in to Hugging Face first if private-repository access is required. For apps, ./start_bonsai_1bit_server.command starts an OpenAI-style API at http://127.0.0.1:8081/v1. Your app must connect to it and supply/execute its tools; the model itself does not include web search or an image generator.

Technical details and credit

Derived from Prism ML's Bonsai 27B 1-bit model. Original layers 6, 12, 18, 24, 30, 36, 42, 48, 54, and 60 are removed from the weight file. config.json selects ritual_model.py, which represents those positions as identity blocks. The retained tensors were copied without retraining or re-quantization. build_ritual_mlx.py records the transformation. The original Apache 2.0 license and notice are included.

Downloads last month
343
Safetensors
Model size
2B params
Tensor type
U32
·
F16
·
MLX
Hardware compatibility
Log In to add your hardware

1-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for VertexAIco/Bonsai-27b-MLX-Fast

Base model

Qwen/Qwen3.6-27B
Finetuned
(3)
this model

Collection including VertexAIco/Bonsai-27b-MLX-Fast