Nanbeige4.2-3B-OptiQ-4bit

Built with mlx-optiq, the MLX-native toolkit to quantize, fine-tune, and serve LLMs locally on Apple Silicon, no PyTorch and no cloud. Try the Lab · All OptiQ quants · Docs

An OptiQ mixed-precision 4-bit quant of Nanbeige/Nanbeige4.2-3B for Apple Silicon.

Nanbeige4.2 is a looped transformer. It has a 22-layer decoder stack, but the stack is run twice with the same shared weights (num_loops: 2), so a 3B-parameter model does the compute of a much deeper one. Each loop keeps its own KV cache. Stock mlx-lm has no nanbeige class, so OptiQ (0.4.6+) ships a vendored, mlx-native port of the architecture that registers itself with mlx-lm on import optiq. It is a reasoning model and opens its answers with a short chain of thought before the final response.

Install

pip install "mlx-optiq>=0.4.6"

The vendored nanbeige architecture ships in the wheel, so this repo loads under stock mlx-lm once optiq is imported. Older OptiQ releases cannot load it.

Usage

import optiq  # registers the nanbeige architecture with mlx-lm
from mlx_lm import load, generate

model, tok = load("mlx-community/Nanbeige4.2-3B-OptiQ-4bit")
prompt = tok.apply_chat_template(
    [{"role": "user", "content": "What is the capital of France?"}],
    tokenize=False, add_generation_prompt=True)
print(generate(model, tok, prompt, max_tokens=512))

optiq serve --model mlx-community/Nanbeige4.2-3B-OptiQ-4bit serves an OpenAI/Anthropic-compatible API. LoRA fine-tuning uses optiq lora train.

The quant

OptiQ measures each layer's sensitivity to quantization and assigns per-layer bit-widths, rather than putting every layer at the same width. Here that gives a mix of 4-bit and 8-bit layers:

Base model Nanbeige/Nanbeige4.2-3B (bf16)
Quantized layers 155 (93 at 4-bit, 62 at 8-bit)
Average precision 5.5 bits/weight
Size on disk 3.15 GB
Group size 64

The full per-layer bit assignment is in optiq_metadata.json.

Links

Downloads last month
-
Safetensors
Model size
0.9B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mlx-community/Nanbeige4.2-3B-OptiQ-4bit

Quantized
(25)
this model