Instructions to use badtheorylabs/Macaw-4bit-MLX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use badtheorylabs/Macaw-4bit-MLX with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("badtheorylabs/Macaw-4bit-MLX") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use badtheorylabs/Macaw-4bit-MLX with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "badtheorylabs/Macaw-4bit-MLX"
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "badtheorylabs/Macaw-4bit-MLX" } ] } } }Run Pi
# Start Pi in your project directory: pi
- OpenClaw new
How to use badtheorylabs/Macaw-4bit-MLX with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "badtheorylabs/Macaw-4bit-MLX"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "badtheorylabs/Macaw-4bit-MLX" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- MLX LM
How to use badtheorylabs/Macaw-4bit-MLX with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "badtheorylabs/Macaw-4bit-MLX"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "badtheorylabs/Macaw-4bit-MLX" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "badtheorylabs/Macaw-4bit-MLX", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use badtheorylabs/Macaw-4bit-MLX with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "badtheorylabs/Macaw-4bit-MLX"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default badtheorylabs/Macaw-4bit-MLX
Run Hermes
hermes
🦜 Macaw · 4-bit MLX
The build that runs on your Mac. 1.5 GB, ~2 GB of memory, about a second per request on an M2.
Ask for something in plain English and it does it — sends the email, finds the file, reads the PDF, tells you what's draining your battery. Nothing leaves the machine.
Install
pip install mlx-lm
mlx_lm.generate --model badtheorylabs/Macaw-4bit-MLX \
--prompt "what's my battery at?"
Or serve it locally:
mlx_lm.server --model badtheorylabs/Macaw-4bit-MLX --port 8138
Then talk to it like any OpenAI-compatible endpoint, passing your tools:
import requests
tools = [{"type": "function", "function": {
"name": "battery_status",
"description": "Report battery percentage and charging state."}}]
r = requests.post("http://127.0.0.1:8138/v1/chat/completions", json={
"model": "badtheorylabs/Macaw-4bit-MLX",
"messages": [{"role": "user", "content": "what's my battery at?"}],
"tools": tools, "temperature": 0.0, "max_tokens": 128})
print(r.json()["choices"][0]["message"]["content"])
# <|tool_call_start|>[battery_status()]<|tool_call_end|>
You run the tool, feed the result back, and it answers in plain language. The full app — 97 macOS tools, parsing, and a confirmation gate on anything destructive — is at github.com/Badtheorylabs/Macaw.
What it's for
"email ada the q3 numbers and tell her i approved"
"why is my mac slow?"
"read ~/Documents/contract.pdf and summarise it"
"take a screenshot, make a folder called Shots, and move it there"
"am i free tomorrow afternoon?"
Mail · Calendar · Files · Notes · Reminders · Music · Safari · Chrome · System settings · Screen reading · Documents · Diagnostics
Specs
| Download | 1.5 GB |
| Memory | ~2 GB running |
| Speed | ~1.2 s per request (M2) · ~40 tok/s decode |
| Context | 128K |
| Quantization | 4-bit affine, group size 64 |
| Requires | Apple Silicon (M1+), macOS 14+ |
Need BF16 for fine-tuning or GPU serving? That's
badtheorylabs/Macaw.
Private by construction
There is no server. The model is on your disk, the tools run through macOS APIs on your machine, and nothing is transmitted. No account, no telemetry, no exceptions — including from us.
License
Derived from LFM2.5-2.6B under the LFM Open License v1.0. Free commercially below $10M annual revenue. App and tooling are MIT.
© 2026 Bad Theory Labs
- Downloads last month
- -
4-bit