The Manchego mascot: a smiling wedge of Manchego cheese in a black beret, waving

Manchego — MLX 4-bit

Manchego v3 converted with mlx_lm at 4 bits (group size 64; 2.4 GB) for Apple silicon. A 4B decision model: state + question + options in, a probability per option out, one forward pass, no generated text. v3 was trained under the SemIf prompt and must be prompted with it (below). This repository's main is v3 (tag v3); Manchego v2.1 stays at tag v2.1. What v3 is, how it was trained, every result and every weakness: the main card.

Rows 80 task rows: accuracy (CE) 200 real-text rows: accuracy (CE)
Manchego v3, MLX 8-bit, SemIf prompt 0.963 (0.136) 0.785 (0.602)
Manchego v3, MLX 4-bit (this model), SemIf prompt 0.887 (0.239) 0.765 (0.715)
Manchego v3, bf16, SemIf prompt 0.963 (0.141) 0.785 (0.602)
Manchego v2.1, MLX 8-bit, its own short prompt (for context) 0.938 (0.180) 0.910 (0.279)
Manchego v2.1, MLX 4-bit, its own short prompt (for context) 0.875 (0.288) 0.900 (0.290)

Accuracy with cross-entropy in brackets, one prompt per forward pass, temperature 1. Four bits cost 5 of the 80 task rows and 4 of the 200 real-text rows against the 8-bit conversion; prefer the 8-bit conversion unless memory decides.

  • The SemIf rows are the ones that describe v3. v3 was trained only under the SemIf prompt, and manchego-serve 0.2 serves it that way. Under v2.1's short prompt, which v3 never trained on, this conversion reads 0.825 (0.369) and 0.755 (0.730); those rows are kept in the main repository's eval/formats.json for completeness, not as a way to use the model.
  • The rows. 80 rows of the project's own task suite (families v3 trains on; none of these rows is in its training file) and 200 human-labelled rows from the development splits of v2.1's real-text corpora: 150 from WANLI, MultiNLI and Bitext, which v3 also trains on (3 of them also occur, with the same label, in its training file), and 50 from banking77, which v3 never trained on.
  • Against v2.1's 4-bit conversion, each under its own prompt: v3 is ahead on the task rows and behind on the real-text rows (0.765 against 0.900), a quarter of which come from a corpus only v2.1 trained on; on the public benchmarks of the main card (bf16) v3 is ahead.
  • Rendering. The v3 rows were scored through manchego-serve 0.2.0's own path (contract semif, yes/no true first), one question per forward pass at temperature 1.0: the MLX builds on Apple silicon, bf16 on an RTX 5090.
  • Serving temperature. No temperature map is bound to these weights: manchego-serve 0.2 serves them at temperature 1. The main repository's map (choice 1.5, noul 0.2, score 1.0; yes/no sharpened on purpose) is bound to the bf16 weights' hash and is refused here.

Use

pip install "mlx-lm==0.31.3" huggingface_hub (tested with mlx-lm 0.31.3). Text-only conversion: the vision tower and the multi-token-prediction head are dropped, so load it with mlx_lm, not mlx-vlm. English. The SemIf prompt handles 2 to 16 options (contract_semif.py, in the main repository).

import importlib.util
import mlx.core as mx
from huggingface_hub import hf_hub_download, snapshot_download
from mlx_lm import load

spec = importlib.util.spec_from_file_location(
    "contract_semif", hf_hub_download("oraculumai/Manchego", "contract_semif.py", revision="v3"))
semif = importlib.util.module_from_spec(spec); spec.loader.exec_module(semif)
model, tok = load(snapshot_download("oraculumai/Manchego-MLX-4bit", revision="v3"))

r = semif.render("WIN a free cruise! Reply YES now.", "Is this message spam?", [("true", None), ("false", None)], "noul")
text = tok.apply_chat_template(r["messages"], tokenize=False, add_generation_prompt=True, enable_thinking=False)
logits = model(mx.array([tok.encode(text, add_special_tokens=False)]))[0, -1]
codes = mx.array([tok.encode(L, add_special_tokens=False)[0] for L in r["letters"]])
print(dict(zip(r["keys"], mx.softmax(logits[codes].astype(mx.float32)).tolist())))   # {"true": P(yes), "false": P(no)}

Or serve it with manchego-serve 0.2 (System One wire contract, offline):

pip install ".[mlx]"                                  # in a manchego-serve v0.2.0 checkout
manchego-serve-download --repo oraculumai/Manchego-MLX-4bit --revision v3 --out ./manchego-v3-mlx4
manchego-serve --backend mlx --model ./manchego-v3-mlx4 --revision v3 --contract semif

17 to 255 options, or an empty state, go through the state-first prompt (contract_v2.py), which v3 never trained on. Scoring several prompts as one padded batch moves probabilities slightly; score one prompt at a time when exact reproducibility matters.

Read the main card first

The most important points, from the main card: on the 231 public JevBench decisions and on Public8, v3 is ahead of v2.1 (our runs, not official JevBench scores; v3 has no official JevBench result yet). On unfamiliar task types it is level with v2.1 and overconfident. It is below v2.1 on fresh out-of-distribution draws of the project's own families. Questions with more than 16 options go through a prompt it never trained on.

Training data attribution and licences (WANLI, MultiNLI, Bitext, GSM8K, MathDial, Spider, ToolACE, When2Call, NVD/CVE/CWE, ProsocialDialog, deepset prompt-injections, the 42 upstream datasets of the Natural Instructions tasks, ViZDoom/Freedoom) are in NOTICE and on the main card, with the declarations the account holder's release decisions rest on. Illustration: the Manchego mascot, an AI-generated image (ChatGPT image generation) supplied by the project's author.

Downloads last month
137
Safetensors
Model size
4B params
Tensor type
U32
·
BF16
·
F32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for oraculumai/Manchego-MLX-4bit

Finetuned
Qwen/Qwen3.5-4B
Quantized
(2)
this model

Datasets used to train oraculumai/Manchego-MLX-4bit