How to use from the
Use from the
MLX library
# Make sure mlx-vlm is installed
# pip install --upgrade mlx-vlm

from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template
from mlx_vlm.utils import load_config

# Load the model
model, processor = load("ToPo-ToPo/Inkling-Small-mlx-2bit")
config = load_config("ToPo-ToPo/Inkling-Small-mlx-2bit")

# Prepare input
image = ["http://images.cocodataset.org/val2017/000000039769.jpg"]
prompt = "Describe this image."

# Apply chat template
formatted_prompt = apply_chat_template(
    processor, config, prompt, num_images=1
)

# Generate output
output = generate(model, processor, formatted_prompt, image)
print(output)

ToPo-ToPo/Inkling-Small-mlx-2bit

MLX 2bit conversion of thinkingmachines/Inkling-Small for Apple Silicon (mlx-vlm). 276B total / 12B active sparse MoE (42 layers, 256 routed experts top-6 + 2 shared), text + image + audio in, text out.

See also: ToPo-ToPo/Inkling-Small-mlx-4bit.

Requires mlx-vlm >= 0.6.9

0.6.9 is the first release whose models/inkling can load an official Inkling checkpoint through the public loader, and the first that implements the MoE global_scale / gate.bias tensors. On 0.6.7 / 0.6.8 this repo will not load.

from mlx_vlm import load, generate
model, processor = load("ToPo-ToPo/Inkling-Small-mlx-2bit")

The config is the official schema, unmodified — no key translation and no loader patches are needed.

Provenance (self-converted from official weights)

  • Source: thinkingmachines/Inkling-Small (license: apache-2.0, bf16, 531.9 GB)
  • Tool: mlx-vlm 0.6.9mlx_vlm.convert --hf-path thinkingmachines/Inkling-Small --mlx-path . -q --q-bits 2 --q-group-size 64
  • Effective: 2.506 bits/weight (77 GiB on disk)
  • Only edit on top of the conversion: pad_token / eos_token added to tokenizer_config.json (the official TokenizersBackend config sets neither, so transformers raises on any padded call). Both point at existing ids — the vocabulary is unchanged.

Reasoning effort

The chat template always injects a Thinking effort level: system message (default 0.9). Control it with the OpenAI-compatible reasoning_effort"none" / "minimal" / "low" / "medium" / "high" / "max", or a float in [0.0, 0.99]. "none" disables thinking entirely.

When serving over mlx_vlm.server, note that Inkling wraps its answer in structural tokens (<|message_model|>, <|content_text|>, <|end_message|>) which the server's fixed _CONTENT_MARKERS list does not strip, and that its reasoning channel is <|content_thinking|><|end_message|><|message_model|> rather than one of the built-in marker pairs. Set MLX_VLM_THINKING_START_TOKEN / MLX_VLM_THINKING_END_TOKEN accordingly and strip the structural tokens, or the reasoning and those markers end up in content.

Revision history

  • 2026-08-04 — reconverted with mlx-vlm 0.6.9. The previous upload had been converted with 0.6.7, whose models/inkling did not implement the MoE mlp.global_scale (50 keys) and mlp.gate.bias (40 keys) present in the official checkpoint, so those tensors were silently dropped. It also shipped a translated config (renamed intermediate_size / dense_intermediate_size, etc.) that 0.6.9 rejects. If you pulled this repo before this date, re-download it.
Downloads last month
225
Safetensors
Model size
25B params
Tensor type
BF16
·
U32
·
F32
·
MLX
Hardware compatibility
Log In to add your hardware

2-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ToPo-ToPo/Inkling-Small-mlx-2bit

Quantized
(42)
this model