Text Generation
MLX
Safetensors
English
qwen3
agent
agentic
tool-use
function-calling
orchestration
magentic
conversational
4-bit precision
Instructions to use mlx-community/MagenticBrain-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use mlx-community/MagenticBrain-4bit with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("mlx-community/MagenticBrain-4bit") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use mlx-community/MagenticBrain-4bit with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "mlx-community/MagenticBrain-4bit"
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "mlx-community/MagenticBrain-4bit" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent new
How to use mlx-community/MagenticBrain-4bit with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "mlx-community/MagenticBrain-4bit"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default mlx-community/MagenticBrain-4bit
Run Hermes
hermes
- OpenClaw new
How to use mlx-community/MagenticBrain-4bit with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "mlx-community/MagenticBrain-4bit"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "mlx-community/MagenticBrain-4bit" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- MLX LM
How to use mlx-community/MagenticBrain-4bit with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "mlx-community/MagenticBrain-4bit"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "mlx-community/MagenticBrain-4bit" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "mlx-community/MagenticBrain-4bit", "messages": [ {"role": "user", "content": "Hello"} ] }'
File size: 5,757 Bytes
0c0f435 131a94c 0c0f435 131a94c 0c0f435 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 | ---
library_name: mlx
license: mit
pipeline_tag: text-generation
base_model: microsoft/MagenticBrain
language:
- en
tags:
- agent
- agentic
- tool-use
- function-calling
- orchestration
- magentic
- qwen3
- mlx
---
# mlx-community/MagenticBrain-4bit
[microsoft/MagenticBrain](https://huggingface.co/microsoft/MagenticBrain) converted to MLX and quantized to
**4-bit**, for inference on Apple Silicon.
MagenticBrain is a 14.8B orchestration model from Microsoft Research AI
Frontiers, supervised fine-tuned from Qwen3-14B for planning, tool selection,
multi-turn tool chaining and sub-agent delegation. It is **not a general-purpose
chat model** — see the [original card](https://huggingface.co/microsoft/MagenticBrain).
## Quantization
| | |
|---|---|
| Requested bits | 4 |
| Group size | 64 |
| Mode | affine |
| **Effective bits per weight** | **4.5** |
| On-disk size | 7.8 GB |
| Shards | 8 |
The source is stored in **float32** (14.8B params × 4 bytes = 59 GB on disk),
which is why the repo is roughly twice the size of a typical bf16 release. Weights
are cast to bf16 before quantizing.
## Fidelity vs the original weights
Measured against the **fp32 source**, streamed tensor-by-tensor from disk, over
all **14,767,882,240** parameters. No prompts or sampling involved — this is a direct
measurement of how much numerical information the quantization discarded.
| Metric | 4-bit |
|---|---|
| Relative L2 error | **9.30%** |
| Cosine similarity | **0.995684** |
| Signal-to-quantization-noise | **20.63 dB** |
| Worst single-element error | 0.098161 |
Both variants, for comparison:
| Variant | bpw | Relative L2 | Cosine | SNR | Size |
|---|---|---|---|---|---|
| 4-bit | 4.5 | 9.30% | 0.995684 | 20.63 dB | 7.8 GB |
| 8-bit | 8.5 | 0.73% | 0.999973 | 42.68 dB | 15 GB |
Highest-error tensors in this variant (the early layers are consistently the
most sensitive):
```
rel_l2=0.11013 snr= 19.16 dB model.layers.1.self_attn.q_proj
rel_l2=0.10583 snr= 19.51 dB model.layers.1.self_attn.k_proj
rel_l2=0.10436 snr= 19.63 dB model.layers.1.self_attn.v_proj
rel_l2=0.10272 snr= 19.77 dB model.layers.1.mlp.gate_proj
rel_l2=0.10245 snr= 19.79 dB model.layers.4.mlp.down_proj
rel_l2=0.09928 snr= 20.06 dB model.layers.3.mlp.down_proj
```
## Why there is no bf16 behavioural control
The methodology used for these conversions compares a quantized model's
*outputs* against the unquantized original. **That was not possible here**, and
the reason is worth stating rather than omitting:
At 14.8B parameters, bf16 weights are ~28 GB. On the 32 GB machine used for this
conversion, loading them drove the system into swap — measured at 34.9 GB of
35.8 GB swap in use, with 50 tokens taking over 10 minutes. Any benchmark run
under those conditions would measure paging, not the model.
So the comparison against the original is done at the **weight level** (above),
which is exact and hardware-independent, and behavioural benchmarks are run on
the variants that actually fit in memory. What is *not* claimed anywhere here is
"indistinguishable from bf16 in behaviour" — that would require a control this
hardware cannot run.
## Tool calling (BFCL)
[Berkeley Function-Calling Leaderboard](https://gorilla.cs.berkeley.edu/leaderboard.html)
v4, scored with **AST checking** against BFCL's ground truth: correct function
selected, all required parameters present, each argument matching BFCL's list of
accepted values, types normalised, no invented parameters. Deterministic — no
judge involved.
| Category | Accuracy | Parse rate | n |
|---|---|---|---|
| `live_simple` | 0.900 | 0.925 | 40 |
| `live_multiple` | 0.675 | 0.925 | 40 |
| `multiple` | 0.775 | 0.925 | 40 |
| `parallel` | 0.625 | 0.825 | 40 |
| **Overall** | **0.744** | | 160 |
Categories: `live_simple` (one function, real user queries), `live_multiple` and
`multiple` (must select among several), `parallel` (several calls in one turn).
Accuracy degrading toward `parallel` is expected — it is the hardest category.
## Throughput
Decode speed and memory on the machine used for conversion (M2 Pro, 32 GB) are
reported in the project notes rather than here, since they do not transfer across
chips. The practical point: at 4-bit the model needs roughly 7.8 GB of
weights, which fits comfortably in 32 GB alongside a working KV cache.
## Usage
```bash
pip install mlx-lm
```
```python
from mlx_lm import load, generate
model, tokenizer = load("mlx-community/MagenticBrain-4bit")
tools = [{
"type": "function",
"function": {
"name": "web_search",
"description": "Search the web",
"parameters": {
"type": "object",
"properties": {"query": {"type": "string"}},
"required": ["query"],
},
},
}]
prompt = tokenizer.apply_chat_template(
[{"role": "user", "content": "Find the 2026 Turing Award winner."}],
tools=tools, tokenize=False, add_generation_prompt=True,
)
print(generate(model, tokenizer, prompt, max_tokens=256, verbose=False))
```
The model emits structured JSON tool calls and selects only from the tools you
declare. Your harness is responsible for parsing the calls, executing them, and
handling the `submit` terminator — see the original card for the protocol.
## What was not measured
No IFEval or general-knowledge benchmarks were run. No agentic end-to-end
evaluation inside MagenticLite (Microsoft's harness, which the model was
co-designed with) was performed. If your use case is the full orchestration loop,
evaluate on your own tasks.
## Credits
All credit for the model belongs to Microsoft Research AI Frontiers. This is a
format conversion and quantization; no training or fine-tuning was performed.
Licensed MIT, as the original.
|