Instructions to use VertexAIco/Bonsai-27b-MLX-Fast with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use VertexAIco/Bonsai-27b-MLX-Fast with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("VertexAIco/Bonsai-27b-MLX-Fast") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use VertexAIco/Bonsai-27b-MLX-Fast with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "VertexAIco/Bonsai-27b-MLX-Fast"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "VertexAIco/Bonsai-27b-MLX-Fast" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use VertexAIco/Bonsai-27b-MLX-Fast with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "VertexAIco/Bonsai-27b-MLX-Fast"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "VertexAIco/Bonsai-27b-MLX-Fast" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "VertexAIco/Bonsai-27b-MLX-Fast", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use VertexAIco/Bonsai-27b-MLX-Fast with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "VertexAIco/Bonsai-27b-MLX-Fast"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default VertexAIco/Bonsai-27b-MLX-Fast
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use VertexAIco/Bonsai-27b-MLX-Fast with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "VertexAIco/Bonsai-27b-MLX-Fast"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "VertexAIco/Bonsai-27b-MLX-Fast" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Bonsai 27b MLX Fast
Benchmark at a glance
M4 Mac mini · 16 GB memory · short prompt · model already loaded.
| Version | Output tokens/sec | Speed vs original | First token | Task checks |
|---|---|---|---|---|
| Original 1-bit Bonsai, stock-style settings | 19.7 | 1.00× | 1.62 s | Not separately checked |
| Optimization | 19.8 | 1.00×; no meaningful gain | 1.60 s | 15/18 |
| Fast | 22.0 | 1.12× | 1.36 s | 11/18 |
| UltraFastUnStable, revised | 24.6 | 1.25× | 1.35 s | 11/18, earlier same-candidate check |
Read this before comparing: these are three-run medians for the same 81-token input and 128 output tokens, with thinking off and a tiny warm-up. Loading time is excluded. The revised UltraFastUnStable was measured later with a different memory ceiling, so its ratio is a comparison of recorded results—not a fresh same-session, controlled speedup. Longer inputs and tool schemas take longer. A higher rate does not mean better or correct answers; task counts are not percentages of intelligence retained.
The retired 48-block experiment reached 79.5 tokens/sec / 0.427 s TTFT, but produced nonsense or blank lines. It is unusable, is not the current UltraFastUnStable download, and is kept only in its historical revision tag. Measurement details and limitations.
Quick tour: modestly faster, less reliable
- What it is: ten original blocks physically removed; 54 remain. Its remaining weights are unchanged. It saves about 600 MB and some generation time, but lost four task passes versus Optimization.
- What to download: the entire repository through Files and versions
or the CLI command below. Pruning is inside the checkpoint, but
ritual_model.pyand Prism's runtime are still required—no speed flag needed. - How to try it: follow Get started with a short, low-stakes request. Check facts and validate tool calls. For strict JSON or dependable instruction following, choose Optimization instead.
A smaller, slightly faster—but less reliable—edition of Bonsai for Mac. This edition removes ten of the original model's 64 layers. Its remaining weights are unchanged. Removing layers saves work and storage, but it also changes the model's answers.
This is an experimental speed/quality tradeoff, not a better model overall. For reliable tool calls, factual answers, strict JSON, and following instructions, choose Bonsai 27b MLX Optimization.
The difference in plain English
| Optimization | Fast — this page | |
|---|---|---|
| Model layers | All 64 kept | Ten removed; 54 kept |
| Weight-file size | About 5.13 GB | About 4.53 GB |
| Measured output speed | 19.8 tokens/second | 22.0 tokens/second |
| Time to first output token | 1.60 seconds | 1.36 seconds |
| Local task check | 15/18 passed | 11/18 passed |
The speed benefit was about 2.3 extra tokens per second and 0.24 second less waiting for the first token. That is a modest improvement, not a huge speed jump. The weight file is about 600 MB smaller, and peak MLX memory was 4.06 GB instead of 4.71 GB.
Read the quality warning before downloading
In our small 18-task check, Fast passed four fewer tasks than Optimization: a 22 percentage-point drop in task pass rate. That is not a measurement of its overall intelligence, but it is a clear warning for the workloads tested.
Failures included a wrong multiplication result, invalid JSON, inventing an 80% success rate that was not present in the supplied information, and refusing to schedule rather than asking for missing details. A recursion explanation also became repetitive. Both editions failed some clarification requests.
Fast has not demonstrated the project's intended low-loss quality tradeoff. The earlier 13% tolerance was not an overall intelligence measurement, and this small task check cannot establish such a percentage. Fast should not be used when dependable facts, valid tool payloads, or strict instructions are important. Later experimental variants reached around 24–26 tokens/second but also failed capability checks, so they were rejected and are not included in this download.
See the detailed test results for the task checklist, timing conditions, and rejected experiments.
What the speed numbers mean
Tests used an M4 Mac mini with 16 GB memory, the same 81-token input, and 128 generated tokens per run. The table uses medians of three runs, with thinking off and greedy decoding. These are actual output-generation rates, not hidden-reasoning counts.
The model was already loaded and warmed up. Loading is not included in time to first token, and longer inputs can take much longer. These numbers are not a promise of a two-second response for every request or every Mac. Peak MLX memory is not the same as total system memory.
Important: one weight file is not enough
Use Files and versions and download the entire repository: weights, configuration, tokenizer,
ritual_model.py, installer, and launch code. A singlemodel.safetensorsfile does not contain the complete optimized setup. An app's “Download MLX” button does not install this required code or runtime for you.
The layer removal is already part of this checkpoint; no extra speed-mode flag is needed. However, the included custom loader is required to use these weights correctly, and Prism ML's MLX fork supplies the 1-bit kernels. This is not a universal drag-and-drop checkpoint for every MLX app. Use only model code you trust.
Get started
You need an Apple Silicon Mac, uv, and Apple's Metal toolchain. The installer
builds the required runtime with one compilation job; compilation still takes
time and space. The compiled runtime is not included.
hf download VertexAIco/Bonsai-27b-MLX-Fast --local-dir Bonsai-27b-MLX-Fast
cd Bonsai-27b-MLX-Fast
./install_prism_mlx.command
.venv-1bit/bin/python run_bonsai_1bit.py --interactive
Sign in to Hugging Face first if private-repository access is required.
For apps, ./start_bonsai_1bit_server.command starts an OpenAI-style API at
http://127.0.0.1:8081/v1. Your app must connect to it and supply/execute its
tools; the model itself does not include web search or an image generator.
Technical details and credit
Derived from Prism ML's Bonsai 27B 1-bit model.
Original layers 6, 12, 18, 24, 30, 36, 42, 48, 54, and 60 are removed from the
weight file. config.json selects ritual_model.py, which represents those
positions as identity blocks. The retained tensors were copied without
retraining or re-quantization. build_ritual_mlx.py records the transformation.
The original Apache 2.0 license and notice are included.
- Downloads last month
- 343
1-bit