Instructions to use VertexAIco/Bonsai-27b-MLX-UltraFastUnStable with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use VertexAIco/Bonsai-27b-MLX-UltraFastUnStable with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("VertexAIco/Bonsai-27b-MLX-UltraFastUnStable") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use VertexAIco/Bonsai-27b-MLX-UltraFastUnStable with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "VertexAIco/Bonsai-27b-MLX-UltraFastUnStable"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "VertexAIco/Bonsai-27b-MLX-UltraFastUnStable" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use VertexAIco/Bonsai-27b-MLX-UltraFastUnStable with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "VertexAIco/Bonsai-27b-MLX-UltraFastUnStable"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "VertexAIco/Bonsai-27b-MLX-UltraFastUnStable" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "VertexAIco/Bonsai-27b-MLX-UltraFastUnStable", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use VertexAIco/Bonsai-27b-MLX-UltraFastUnStable with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "VertexAIco/Bonsai-27b-MLX-UltraFastUnStable"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default VertexAIco/Bonsai-27b-MLX-UltraFastUnStable
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use VertexAIco/Bonsai-27b-MLX-UltraFastUnStable with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "VertexAIco/Bonsai-27b-MLX-UltraFastUnStable"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "VertexAIco/Bonsai-27b-MLX-UltraFastUnStable" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
UltraFastUnStable
Benchmark at a glance
M4 Mac mini · 16 GB memory · short prompt · model already loaded.
| Version | Output tokens/sec | Speed vs original | First token | Task checks |
|---|---|---|---|---|
| Original 1-bit Bonsai, stock-style settings | 19.7 | 1.00× | 1.62 s | Not separately checked |
| Optimization | 19.8 | 1.00×; no meaningful gain | 1.60 s | 15/18 |
| Fast | 22.0 | 1.12× | 1.36 s | 11/18 |
| UltraFastUnStable, revised | 24.6 | 1.25× | 1.35 s | 11/18, earlier same-candidate check |
Read this before comparing: these are three-run medians for the same 81-token input and 128 output tokens, with thinking off and a tiny warm-up. Loading time is excluded. The revised UltraFastUnStable was measured later with a different memory ceiling, so its ratio is a comparison of recorded results—not a fresh same-session, controlled speedup. Longer inputs and tool schemas take longer. A higher rate does not mean better or correct answers; task counts are not percentages of intelligence retained.
The retired 48-block experiment reached 79.5 tokens/sec / 0.427 s TTFT, but produced nonsense or blank lines. It is unusable, is not the current UltraFastUnStable download, and is kept only in its historical revision tag. Measurement details and limitations.
Quick tour: faster experiment that still makes mistakes
- What it is: 54 full blocks plus tiny corrections replacing ten removed blocks. Both pruning and correction vectors are embedded in the main checkpoint. This is the partially usable revision, not the broken 48-block one.
- What to download: the entire repository through Files and versions or the CLI command below. The custom loader and Prism runtime remain necessary; no separate adapter or speed-mode flag is needed for inference.
- How to try it: follow Get started. Use low-stakes prompts, inspect every important answer, and validate tool calls. It can produce useful replies, but math, clarification, and strict formatting remain unreliable.
An experimental faster Bonsai that can still produce useful answers—but makes more mistakes. This is the revised, less-aggressive edition, not the original 48-block version that produced nonsense.
What changed?
Ten of the original 64 decoder blocks are replaced with tiny correction vectors, leaving 54 full decoder blocks. Their original attention and feed-forward weights are physically removed from this checkpoint. The small corrections approximate the missing blocks; they do not restore the original intelligence. Retained original weights are unchanged.
The pruning and correction vectors are inside model.safetensors. You do
not need to enable a speed flag or provide a separate adapter to run it.
Is it usable?
Partially, for low-stakes experiments—not as a dependable assistant. The same candidate passed 11 of 18 earlier task checks: it could request web search, construct image-tool calls, explain a Python error, summarize museum information, and avoid inventing a success rate. It failed other arithmetic, strict JSON-only formatting, clarification, and some instruction checks.
Its explanation of recursion was brief and incomplete, not a strong teaching answer. Its transit summary omitted an alternative route. It sometimes fills in missing design details instead of asking. Check every important answer and validate tool requests before executing them.
That small task check is not a percentage of intelligence retained. This edition deliberately trades reliability for speed; no low-loss guarantee is claimed. Choose Optimization for the full original model.
Speed and memory
Measured on an M4 Mac mini with 16 GB unified memory:
| Metric | Revised edition |
|---|---|
| Output generation, median of three runs | 24.6 tokens/second |
| Time to first token, excluding loading | 1.35 seconds |
| Peak MLX memory | 4.06 GB |
| Weight file size | 4.53 GB |
Earlier measurements were 22.0 tokens/sec for Fast and 19.8 for Optimization. Those were not rerun in this session, so this is not a fresh paired comparison. The speed-test summary was coherent but included unsupported assumptions and hit its output cap; this rate is not a guarantee of correct completed answers.
See the benchmark report for fresh measurements, conditions, failures, and the earlier 18-task result. Speed is actual output tokens, not hidden reasoning. TTFT excludes loading and uses a two-token warm-up. Longer prompts, tool schemas, and different Macs can take much longer.
This is not the 79.5 tokens/sec model. That earlier 48-block edition collapsed
into nonsense and is preserved only under the extreme-48-blocks revision tag.
Its speed must not be attributed to this revised checkpoint.
Download the whole repository
Open Files and versions and download the entire repository, including weights, config, tokenizer,
unstable_model.py,ritual_model.py, installer, and launch code. One safetensors file is not a complete runnable setup. A “Download MLX” button that fetches only weights may omit required code.
The custom loader reads the embedded corrections, and Prism ML's 1-bit MLX fork supplies the kernel. This is not a universal drop-in checkpoint for every MLX application. Use only model code you trust.
Get started on Apple Silicon
You need uv and Apple's Metal toolchain. The installer compiles the runtime
with one job and clears temporary build files; allow extra space for compilation.
The compiled runtime is not shipped.
hf download VertexAIco/Bonsai-27b-MLX-UltraFastUnStable --local-dir UltraFastUnStable
cd UltraFastUnStable
./install_prism_mlx.command
.venv-1bit/bin/python run_bonsai_1bit.py "What is 17 times 23?" --max-tokens 64
Sign in to Hugging Face if private access is required. The server launcher offers an OpenAI-style API; your app must supply and execute tools. The model has no built-in web browser or image generator. Keep initial requests short.
Technical notes and credit
Original blocks 4, 5, 8, 9, 10, 12, 13, 14, 16, 17 are replaced by per-channel
x * scale + bias operations. Corrections were fitted previously against this
exact source checkpoint on six short calibration cases (238 tokens); this is
a tiny calibration, not retraining. No speculative decoding or additional
quantization is used. The original 64-slot cache layout is preserved.
The small onebit-affine-v1.safetensors file is included for reproducible builds
only; inference uses corrections embedded in the main checkpoint. Rebuilding
requires the original source model, which is not duplicated in this repository.
Derived from Prism ML Bonsai 27B 1-bit. Original license and notices are included. Fast and Optimization are unchanged.
- Downloads last month
- 336
1-bit