UltraFastUnStable

Benchmark at a glance

M4 Mac mini · 16 GB memory · short prompt · model already loaded.

Version Output tokens/sec Speed vs original First token Task checks
Original 1-bit Bonsai, stock-style settings 19.7 1.00× 1.62 s Not separately checked
Optimization 19.8 1.00×; no meaningful gain 1.60 s 15/18
Fast 22.0 1.12× 1.36 s 11/18
UltraFastUnStable, revised 24.6 1.25× 1.35 s 11/18, earlier same-candidate check

Read this before comparing: these are three-run medians for the same 81-token input and 128 output tokens, with thinking off and a tiny warm-up. Loading time is excluded. The revised UltraFastUnStable was measured later with a different memory ceiling, so its ratio is a comparison of recorded results—not a fresh same-session, controlled speedup. Longer inputs and tool schemas take longer. A higher rate does not mean better or correct answers; task counts are not percentages of intelligence retained.

The retired 48-block experiment reached 79.5 tokens/sec / 0.427 s TTFT, but produced nonsense or blank lines. It is unusable, is not the current UltraFastUnStable download, and is kept only in its historical revision tag. Measurement details and limitations.

Quick tour: faster experiment that still makes mistakes

  1. What it is: 54 full blocks plus tiny corrections replacing ten removed blocks. Both pruning and correction vectors are embedded in the main checkpoint. This is the partially usable revision, not the broken 48-block one.
  2. What to download: the entire repository through Files and versions or the CLI command below. The custom loader and Prism runtime remain necessary; no separate adapter or speed-mode flag is needed for inference.
  3. How to try it: follow Get started. Use low-stakes prompts, inspect every important answer, and validate tool calls. It can produce useful replies, but math, clarification, and strict formatting remain unreliable.

An experimental faster Bonsai that can still produce useful answers—but makes more mistakes. This is the revised, less-aggressive edition, not the original 48-block version that produced nonsense.

What changed?

Ten of the original 64 decoder blocks are replaced with tiny correction vectors, leaving 54 full decoder blocks. Their original attention and feed-forward weights are physically removed from this checkpoint. The small corrections approximate the missing blocks; they do not restore the original intelligence. Retained original weights are unchanged.

The pruning and correction vectors are inside model.safetensors. You do not need to enable a speed flag or provide a separate adapter to run it.

Is it usable?

Partially, for low-stakes experiments—not as a dependable assistant. The same candidate passed 11 of 18 earlier task checks: it could request web search, construct image-tool calls, explain a Python error, summarize museum information, and avoid inventing a success rate. It failed other arithmetic, strict JSON-only formatting, clarification, and some instruction checks.

Its explanation of recursion was brief and incomplete, not a strong teaching answer. Its transit summary omitted an alternative route. It sometimes fills in missing design details instead of asking. Check every important answer and validate tool requests before executing them.

That small task check is not a percentage of intelligence retained. This edition deliberately trades reliability for speed; no low-loss guarantee is claimed. Choose Optimization for the full original model.

Speed and memory

Measured on an M4 Mac mini with 16 GB unified memory:

Metric Revised edition
Output generation, median of three runs 24.6 tokens/second
Time to first token, excluding loading 1.35 seconds
Peak MLX memory 4.06 GB
Weight file size 4.53 GB

Earlier measurements were 22.0 tokens/sec for Fast and 19.8 for Optimization. Those were not rerun in this session, so this is not a fresh paired comparison. The speed-test summary was coherent but included unsupported assumptions and hit its output cap; this rate is not a guarantee of correct completed answers.

See the benchmark report for fresh measurements, conditions, failures, and the earlier 18-task result. Speed is actual output tokens, not hidden reasoning. TTFT excludes loading and uses a two-token warm-up. Longer prompts, tool schemas, and different Macs can take much longer.

This is not the 79.5 tokens/sec model. That earlier 48-block edition collapsed into nonsense and is preserved only under the extreme-48-blocks revision tag. Its speed must not be attributed to this revised checkpoint.

Download the whole repository

Open Files and versions and download the entire repository, including weights, config, tokenizer, unstable_model.py, ritual_model.py, installer, and launch code. One safetensors file is not a complete runnable setup. A “Download MLX” button that fetches only weights may omit required code.

The custom loader reads the embedded corrections, and Prism ML's 1-bit MLX fork supplies the kernel. This is not a universal drop-in checkpoint for every MLX application. Use only model code you trust.

Get started on Apple Silicon

You need uv and Apple's Metal toolchain. The installer compiles the runtime with one job and clears temporary build files; allow extra space for compilation. The compiled runtime is not shipped.

hf download VertexAIco/Bonsai-27b-MLX-UltraFastUnStable --local-dir UltraFastUnStable
cd UltraFastUnStable
./install_prism_mlx.command
.venv-1bit/bin/python run_bonsai_1bit.py "What is 17 times 23?" --max-tokens 64

Sign in to Hugging Face if private access is required. The server launcher offers an OpenAI-style API; your app must supply and execute tools. The model has no built-in web browser or image generator. Keep initial requests short.

Technical notes and credit

Original blocks 4, 5, 8, 9, 10, 12, 13, 14, 16, 17 are replaced by per-channel x * scale + bias operations. Corrections were fitted previously against this exact source checkpoint on six short calibration cases (238 tokens); this is a tiny calibration, not retraining. No speculative decoding or additional quantization is used. The original 64-slot cache layout is preserved.

The small onebit-affine-v1.safetensors file is included for reproducible builds only; inference uses corrections embedded in the main checkpoint. Rebuilding requires the original source model, which is not duplicated in this repository.

Derived from Prism ML Bonsai 27B 1-bit. Original license and notices are included. Fast and Optimization are unchanged.

Downloads last month
336
Safetensors
Model size
2B params
Tensor type
U32
·
F16
·
MLX
Hardware compatibility
Log In to add your hardware

1-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for VertexAIco/Bonsai-27b-MLX-UltraFastUnStable

Base model

Qwen/Qwen3.6-27B
Finetuned
(3)
this model

Collection including VertexAIco/Bonsai-27b-MLX-UltraFastUnStable