The Manchego mascot: a smiling wedge of Manchego cheese in a black beret, waving

Manchego

A 4B model for typed decisions. One forward pass, a probability for every option.

Give Manchego a state (text or JSON), a question and a closed set of options. It returns a probability for each one: pick an option (choice), a yes/no condition (noul), or an ordered level (score). It does not generate text. This is Manchego v3: Qwen3.5-4B with a LoRA adapter (merged here), trained from the untrained base under the SemIf prompt on about 100M tokens of decision rows. Manchego v2.1, the previous release, stays available at tag v2.1.

What changed in v3

  • Trained from the base, not continued from v2.1. A fresh LoRA (rank 16) on the attention and gated-delta-net projections of every layer.
  • A different prompt. SemIf's direct-options prompt: a system message and one JSON user message holding the evidence, the criterion and lettered options. Serve v3 with it ("contract": "semif", manchego-serve 0.2); v2.1's short prompt is not v3's.
  • More data. Hard computed families, answer judging, eight more public corpora (math, SQL, tool calling, security, dialogue safety, prompt injection) and a self-distilled replay toward the base.
  • A serving temperature map that sharpens yes/no answers on purpose (below).

Results

Suite (our runs, not official JevBench scores) rows untrained Qwen3.5-4B, v2.1's prompt untrained Qwen3.5-4B, SemIf prompt Manchego v2.1 Manchego v3 v3 minus v2.1 [95% interval]
JevBench, 231 public decisions (our run, not an official JevBench score) 231 0.775 (0.493) 0.801 (0.485) 0.805 (0.423) 0.857 (0.326) +0.052 [+0.004, +0.102]
Public8, eight public real-text datasets 1,600 0.686 (0.843) 0.725 (0.847) 0.732 (0.684) 0.753 (0.609) +0.021 [+0.002, +0.039]
27 held-out Natural Instructions tasks (source-clean) 1,561 0.678 (0.919) 0.678 (1.056) 0.701 (0.929) 0.709 (0.941) +0.007 [-0.016, +0.031]

Accuracy with cross-entropy in brackets, temperature 1, one runtime for every model (see How the comparison was run). The JevBench rows on this card are our runs on its 231 public decisions, not official JevBench scores. JevBench's official score also reads sealed items, calibration, speed and cost, and only its maintainer runs it. v3 has no official JevBench result yet.

  • JevBench's public decisions (our run, not an official JevBench score): v3 is ahead of v2.1 (+0.052, interval excludes zero; cross-entropy -0.097 [-0.171, -0.032]) and of the untrained base under v2.1's prompt (+0.082 [+0.022, +0.143]). The untrained base itself reads 0.801 under the SemIf prompt, so part of that step is the prompt.
  • Public8: ahead of v2.1 (+0.021, resolved) with better probabilities (cross-entropy -0.075 [-0.107, -0.044]).
  • Held-out task types: level with v2.1 (+0.007, interval through zero) and +0.031 [+0.002, +0.065] above the untrained base. Its probabilities there are no better than the base's (cross-entropy 0.941 against 0.919) and are overconfident (see Calibration).

Sealed sets, read once. Two sets held back for this read, with no recorded Manchego training or scoring on either, were opened once, on 2026-09-30, after v3 was fixed; nothing was selected by that read. The unseen-task set's record also notes earlier exposure: before it was reserved, Jev (TypeSafe AI's hosted decision service) was run on training rows of its tasks, and researchers saw some of its task names.

Sealed set, read once rows untrained Qwen3.5-4B, v2.1's prompt untrained Qwen3.5-4B, SemIf prompt Manchego v2.1 Manchego v3 v3 minus v2.1 [95% interval]
24 unseen Natural Instructions tasks (13 source clusters) 5,397 0.641 (0.918) 0.674 (0.911) 0.659 (0.849) 0.695 (0.861) +0.042 [-0.018, +0.107]
Fresh out-of-distribution draws of four of the project's own computed families 1,193 0.445 (1.161) 0.508 (1.091) 0.795 (0.466) 0.762 (0.578) -0.033 [-0.060, -0.005]

The model columns pool rows. The unseen-task contrast is the mean over the 24 tasks of the per-task difference, with whole source clusters resampled; v3 is also +0.060 [-0.020, +0.143] above the untrained base there, again not resolved. The family draws (policy decisions, lookup chains, temporal reasoning, answer adequacy; 400 instances resampled) are where v2.1 kept training on v2's families and v3 started over from the base: v3 is 0.033 below v2.1 (resolved) and +0.317 [+0.280, +0.353] above the base.

The previous release's official JevBench result

JevBench v1.5.4, official (maintainer-run) score (95% interval) rank Intelligence Calibration Speed Cost
Manchego v2.1 (the previous release, tag v2.1) 68.8 (59.6 to 70.3) 11 of 106 51.2 84.9 88.6 64.3

Measured by JevBench's maintainer on their own offline GPU (RTX6000, bf16, temperature 1.0, original option order) with manchego-serve v0.1.1 and oraculumai/Manchego at the v2.1 commit 77403228; the headline score weighs the four axes equally. This row is v2.1's, not v3's. v3 has no official result yet.

Quick start

v3 is served by manchego-serve 0.2 (Apache-2.0). The server speaks TypeSafe AI's System One wire contract (POST /v1/systemone), runs offline, and hashes the weights it loads and reports whether they are the published ones. Pin this repository by its tag v3 (manchego-serve 0.2.0 pins the commit behind it).

git clone https://github.com/nschlaepfer/manchego-serve && cd manchego-serve && git checkout v0.2.0

# Docker, Linux + NVIDIA GPU: the build downloads the pinned weights once; the container runs offline
docker build -t manchego-serve:3-cuda --build-arg MODEL_REVISION=v3 .
docker run --rm --gpus all -p 127.0.0.1:8000:8000 manchego-serve:3-cuda \
  --contract semif --temperature-map /models/manchego/temperature_map.json

# or pip (install the CUDA build of torch first on Linux + CUDA)
pip install .
manchego-serve-download --repo oraculumai/Manchego --revision v3 --out ./manchego-v3   # uses the network once
manchego-serve --model ./manchego-v3 --revision v3 --contract semif \
  --temperature-map ./manchego-v3/temperature_map.json                                          # offline from here on
curl -s localhost:8000/v1/systemone -H 'content-type: application/json' -d '{
  "model": "manchego-3",
  "state": "Customer: the blender I bought last week smells of burning and stopped working. Order 5521.",
  "questions": {
    "route":     {"type": "choice", "instructions": "Which team should handle this message?",
                  "criteria": {"returns": "refunds, exchanges and defective items",
                               "shipping": "delivery status and lost parcels", "billing": null}},
    "defective": {"type": "noul", "instructions": "Does the customer report a defective product?"},
    "urgency":   {"type": "score", "instructions": "How urgent is this message?",
                  "criteria": ["routine", "soon", "immediately"]}}}'

answers.route.probabilities holds one probability per option, answers.defective.noul is P(yes) and answers.urgency.score is the expected level. Every response names the prompt each question got (manchego.contract_by_question), the temperatures used and the weights' hash. manchego_config.json in this repository already selects semif; --contract semif states it. Use --temperature-map off in place of the map for probabilities at temperature 1, the setting of every number on this card.

Without the server (Transformers). Not a pipeline("text-classification") model: read the option-letter logits. contract_semif.py in this repository renders the prompt v3 was trained on (2 to 16 options; the server handles 17 to 255 with the state-first prompt of contract_v2.py).

# pip install "transformers==5.17.0" torch huggingface_hub      (tested: transformers 5.17.0, torch 2.10.0)
import importlib.util, torch
from huggingface_hub import hf_hub_download
from transformers import AutoModelForCausalLM, AutoTokenizer

REPO, REV = "oraculumai/Manchego", "v3"
spec = importlib.util.spec_from_file_location("contract_semif", hf_hub_download(REPO, "contract_semif.py", revision=REV))
semif = importlib.util.module_from_spec(spec); spec.loader.exec_module(semif)
tok = AutoTokenizer.from_pretrained(REPO, revision=REV)
model = AutoModelForCausalLM.from_pretrained(REPO, revision=REV, dtype=torch.bfloat16, device_map="auto").eval()

def decide(state, question, options, kind):    # options: [(value, description or None)]; noul: [("true", None), ("false", None)]
    r = semif.render(state, question, options, kind)
    text = tok.apply_chat_template(r["messages"], tokenize=False, add_generation_prompt=True, enable_thinking=False)
    ids = tok(text, return_tensors="pt", add_special_tokens=False).to(model.device)
    with torch.no_grad():
        logits = model(**ids).logits[0, -1].float()
    z = logits[[tok.encode(L, add_special_tokens=False)[0] for L in r["letters"]]]
    return dict(zip(r["keys"], torch.softmax(z, 0).tolist()))     # divide z by the type's temperature to match the map

print(decide("WIN a free cruise! Reply YES now.", "Is this message spam?", [("true", None), ("false", None)], "noul"))

Apple silicon: Manchego-MLX-8bit and Manchego-MLX-4bit.

Serving temperatures

Question type serving temperature (temperature_map.json) --temperature-map off
choice 1.5 1.0
noul 0.2 1.0
score 1.0 1.0

The probabilities are softmax(z / T) over the offered option codes, with T set by the question type. A temperature never changes the chosen option.

  • The noul temperature 0.2 sharpens yes/no probabilities on purpose. JevBench v1.5 counts a yes/no answer whose P(yes) lies between 0.20 and 0.80 as wrong, so the map pushes answers out of that band. Under the map, P(yes) is therefore not a calibrated probability wherever the model is unsure. The choice temperature 1.5 flattens choice probabilities; score is unchanged.

  • Every number on this card is at temperature 1. --temperature-map off serves exactly that.

  • How it was chosen (TMAP-V15). Per question type, the temperature from a grid (0.2 to 2.0) that maximises an estimate of JevBench v1.5's composite, computed with v1.5's published scoring rules, on v3's records of development rows it never trained on. No JevBench item chose it:

    TMAP-V15 fit choice noul score total
    development rows 10,485 16,050 1,647 28,182

    The rule was registered before the fit, but after we had seen a noul-temperature sweep on those records and a report on JevBench's public items. After the fit, its effect on the public items was reported and changed nothing.

  • Which weights. The map is bound by hash to the three v3 builds: this repository's bf16 weights (weights_sha256 in Files), where it was fitted, and the MLX 8-bit and 4-bit builds, where it is applied as-is. manchego-serve 0.2 ships it as the default map for those hashes only; any other weights get temperature 1. It was fitted on records scored with the adapter applied to the base in padded batches, not on these merged weights' one-prompt logits; the merge moves option logits by at most 0.125 on its check rows.

Speed

v3's manchego_config.json turns on one piece of manchego-serve 0.2's CUDA fast path by default: the fast host path, which is bit-for-bit the reference arithmetic with less host work (torch backend on CUDA only; --no-fast-host turns it off). CUDA graphs stay off by default. On Linux, which is what the Docker recipe runs, padding each prompt to a graph size made every prompt length slower:

median latency, one decision per request (A10, Linux, Docker image, these weights, SemIf prompt) reference CUDA graphs + fast host
up to 256 prompt tokens 54 ms 71 ms
up to 512 134 ms 143 ms
up to 1,024 190 ms 275 ms
up to 2,048 357 ms 543 ms

On Windows (an RTX 5090) the same graphs were about five times faster than the reference (median 26 vs 124 ms), because per-call overhead dominates there; --cuda-graphs turns them on if that is your platform. They pad each prompt to 256, 512, 1,024 or 2,048 tokens and replay a captured forward: on 83 invented test questions no chosen option changed, but probabilities moved by up to 0.0385 (0.0626 across other bucket sets). Details: manchego-serve's docs/FAST_PATH.md. The official Speed of 88.6 above is v2.1's, measured on the reference path.

Intended use and limits

  • For: typed decisions over options your software supplies (routing, policy checks, triage, graded judgments, judging whether an answer is right), where you read the probabilities.
  • Not for: unreviewed high-stakes decisions (medical, legal, financial, safety) without a person in the loop; open-ended text generation; images (inherited from the base and not measured).
  • Options. v3 trained on 2 to 16 options under SemIf. Questions with 17 to 255 options, or with an empty state, go to the state-first prompt of contract_v2.py, which v3 never trained on. On the 104 rows of the held-out task set with more than 16 options, v3 and v2.1 each got 78 right.
  • Length. v3's longest training prompt was 6,386 tokens. Longer inputs are untested (manchego-serve accepts up to 32,768).
  • Language: English. State: SemIf places the state in a JSON string field; v3's robustness to adversarial state was not measured.
  • Machine-readable: manchego_config.json.

Read before relying on it

  • JevBench-informed. v2's computed families, which v3 retrains on, were designed from an earlier version's per-family JevBench hard-tier scores on the public items. v3's hard families were designed from JevBench's published hard-tier specification and family list, two of them also from this line's per-family public hard-tier results. No JevBench item text was opened, trained on or used to select rows.
  • Public aggregate results. Development used JevBench's published aggregate results, never its items or per-item records: they ordered which new domain generators and public corpora were built first (they set no dose, label or selection), and v1.5's published scoring rules, with v2.1's official row, set the serving temperatures' objective (above).
  • Phrase audit (inherited from v2). An early audit used Jev (TypeSafe AI's hosted decision service) to flag phrases in two template pools of v2's families; the flagged phrases were dropped. Jev never produced a label or a selection for v3.
  • Labels. Targets are computed by our code, or a public dataset's original annotation, or (the self-distilled replay, 15% of the tokens) the untrained base model's own probabilities. ToolACE's and When2Call's labels were generated by their authors' model pipelines. No other label came from a language model, and no hosted decision service or teacher model of ours produced any.
  • Text written by other companies' models. ProsocialDialog (utterances written by GPT-3), ToolACE (dialogues from unnamed generator models), When2Call (NVIDIA's generation pipeline) and MathDial (student turns written by gpt-3.5-turbo) contain model-written text. Their licences allow training and release with attribution; the account holder decided to release v3 with all four.
  • MathDial's reproduction probe. On MathDial, v3 fails the project's trained-text reproduction rule (details under Provenance). The effect is larger on dialogues it never trained on, which points to the dialogues' style rather than memorised rows; the account holder accepted it for this release.
  • Selection. v3 is the final update of its only run; nothing was selected inside it. Earlier candidates for v3 (among them a 300M-token run and an average of two adapters) did not pass their registered development reads. The sealed and public reads above came after, once, and selected nothing.

Training

Base Qwen/Qwen3.5-4B @ 851bf6e8 (hybrid gated-delta-net + attention), untrained
Method LoRA, rank 16, alpha 32, on q/k/v/o of the 8 attention layers and the five in/out projections of the 24 gated-delta-net layers (152 modules, 14,376,960 trainable parameters), merged into the base for release
Objective cross-entropy on the offered option letters' logits at the last prompt position, with exact soft targets where a row declares a distribution; no loss on any text token
Prompt SemIf's direct-options prompt on every row (2 to 16 options); choice options shuffled at every presentation, noul true then false
Run 255,991 rows in one pass, 8,771 updates (token-budget batches of 12,288 padded tokens, at most 64 rows), 99,362,993 prompt tokens; lr 3e-5, 85 warm-up updates, cosine to 0, AdamW, no weight decay, gradient clip 1; one NVIDIA GH200, 2026-09-28, 00:56 to 04:48 UTC; every planned row consumed
Checkpoint the final update; nothing was selected inside the run
Training data (bucket) share of the estimated prompt tokens rows labels
Computed decision families of v2 (policies, lookups, temporal, adequacy, routing, skills, arena and Doom states, score levels), regenerated 27.3% 48,914 computed by our generators
Hard computed families (long policy documents, multi-hop lookups, abstention, temporal strata, exact probabilities, paraphrase invariance, policy compliance) 22.7% 19,214 computed
Answer judging (judging worked answers; judge-error repair) 8.0% 22,547 computed
Human-annotated text: WANLI, MultiNLI, Bitext; 101 Natural Instructions tasks; GSM8K, MathDial, Spider, ToolACE, When2Call, NVD/CVE/CWE, ProsocialDialog, deepset prompt-injections 20.0% 83,701 the datasets' original annotation (ToolACE and When2Call: their authors' generated labels)
Self-distilled replay (a partition of the same public corpora, and fresh answer-judging instances) 15.0% 61,739 the untrained Qwen3.5-4B's own probabilities over the options
Domain generators (auth-log triage, code output, exact draws, genetic crosses, limitation deadlines, loan amortisation, math-answer grading, plan cost sharing) 5.0% 11,087 computed
Calibration rows 2.0% 8,789 computed
Total 100% 255,991

By target, 178,334 rows train toward one gold option, 15,918 toward an exact distribution and 61,739 toward the base's own probabilities. By type: 158,909 choice, 84,115 noul and 12,967 score rows.

No row comes from a benchmark on this card. Every source went through an 8-word overlap gate against the evaluation sets (JevBench's public items, Public8, the held-out and sealed task sets) before the training file was built, and the rows that matched were dropped. Public8's eight datasets and the held-out tasks' upstream datasets appear nowhere in the declared training data (checked by name); what the base model saw in pretraining is unknown. GSM8K and When2Call are also Decision Index benchmarks: v3 trained on their train splits only, and their test splits were kept out of training. "Computed" means exact under the task definition as implemented.

Calibration

Calibration error (ECE, temperature 1) untrained Qwen3.5-4B, v2.1's prompt untrained Qwen3.5-4B, SemIf prompt Manchego v2.1 Manchego v3 v3 mean confidence v3 accuracy
JevBench, 231 public decisions (our run, not an official JevBench score) 0.062 0.059 0.072 0.055 0.854 0.857
Public8 0.122 0.122 0.053 0.055 0.809 0.753
27 held-out Natural Instructions tasks 0.107 0.119 0.083 0.129 0.829 0.709
24 unseen Natural Instructions tasks (sealed) 0.151 0.148 0.098 0.141 0.836 0.695
Fresh out-of-distribution draws of the project's families (sealed) 0.141 0.075 0.051 0.050 0.800 0.762

On familiar ground v3 is as well calibrated as v2.1. On unfamiliar task definitions it is overconfident: on the 27 held-out tasks its mean confidence is 0.829 against an accuracy of 0.709, and its calibration error (0.129) is worse than v2.1's (0.083) and the base's. v2.1's second stage had repaired exactly this in v2; v3, trained from the base, has it again. The serving map's choice temperature of 1.5 flattens choice probabilities; its effect on these sets was not measured.

Weaknesses

  1. Not more accurate than v2.1 on unfamiliar task types (+0.007 on 27 held-out tasks, +0.042 on 24 sealed unseen tasks, both through zero), and overconfident there (above).
  2. Below v2.1 on the project's own families out of distribution (-0.033, resolved, on sealed fresh draws), and on 200 human-labelled rows from the development splits of v2.1's real-text corpora (0.785 against v2.1's 0.910, MLX 8-bit; see the MLX cards). A quarter of those rows come from banking77, which v3 did not train on; v2.1 fitted v2's families and those corpora more closely.
  3. Large menus go through a prompt v3 never trained on (17 to 255 options, or an empty state).
  4. Yes/no probabilities under the serving map are sharp by design, not calibrated. Use --temperature-map off when you need P(yes) as a probability.
  5. JevBench's public items may not predict its sealed items. The step on the 231 public decisions is our run; the official score also reads sealed items, and the project's own instruments have not predicted JevBench's Intelligence axis.
  6. One run, one seed. No second seed of v3 was trained.
  7. Date and number arithmetic in one pass is unreliable outside its own templates. Do the arithmetic in code and ask the model the judgment.
  8. Numerics of serving. Scoring several prompts in one padded batch moves probabilities slightly at bf16 or 8 bits; score one prompt at a time (manchego-serve's default) when exact reproducibility matters.
  9. A 4B model. Its world knowledge is the base model's; it will be wrong on questions that need facts it does not have.

How the comparison was run

Every model in the tables above ran in one runtime on one RTX 5090: bf16, the adapters applied to the base (not merged), padded batches of up to 8,192 tokens, identical rows, temperature 1. v3 ran under the SemIf prompt (the state-first prompt beyond 16 options, as manchego-serve 0.2 serves it); v2.1 and the untrained base ran under v2.1's served policy (the short prompt up to 26 options, state-first beyond), and the base also under SemIf. v2.1 and the untrained base reproduced their archived records on the three public suites (the runtime check passed). Intervals are paired 95% bootstraps: JevBench resamples its scenario groups, Public8 its items within each dataset, the held-out set whole tasks. The merged weights in this repository reproduce base plus adapter within the merge check (merge_record.json: at most 0.125 in any option logit on 12 rows, no changed decision); a bf16 merge is not bit-identical. Aggregates: eval/.

Provenance

  • Adapter: update 8,771, the final update of a fresh LoRA trained from the untrained Qwen3.5-4B; adapter_model.safetensors sha256 c43e2688… (full hash in merge_record.json). Training file sha256 240b5cbc…, 255,991 rows, checked row by row on the training machine by a launch gate: no reserved evaluation task; no target from an external teacher model or a hosted decision service (the replay's targets are the base model's own); every row's permission recorded.
  • A release check walks the permission record of every training row. It cleared v3 on 2026-09-30, after the account holder's decisions above; NOTICE carries every attribution those permissions require.
  • Reproduction probe (verbatim continuation of training texts under teacher forcing, the adapter against the base, texts trained on against held-out texts; eval/de_minimis_probe.json): Bitext, MultiNLI fiction and Spider pass every rule. MathDial fails two: under the decision prompt, 0.075 of trained dialogues are continued verbatim for 8 tokens or more against the base's 0.045 (the rule allows 0.02 more), and 6 trained dialogues are continued for 12 tokens or more by the adapter alone (the rule allows none; as raw text, 2). On 200 MathDial dialogues it never trained on, the same measures read 0.155 against 0.085, and 14.
  • Weights: model.safetensors-* in this repository are the base checkpoint with every language-model tensor replaced by its merged value; the vision tower, the multi-token-prediction head, the config and the tokenizer are the base model's. Manchego was trained and evaluated on text only.

Versions

This repository's main holds Manchego v3 (2026-09-30), tagged v3. Manchego v2.1 (2026-09-21) stays available, unchanged, at tag v2.1 (revision="v2.1"): its weights, card, DETAILS.md and evidence. v2.1 needs its own prompts (manchego-serve's contract auto); v3 needs SemIf. Earlier versions are not published.

Files

Format repository size weights_sha256 (as manchego-serve reports it)
bf16 (Transformers) oraculumai/Manchego 9.3 GB 2ee838433bfe278a226dc644667ad4a99ece82cc47325c7645a7dae723c1863b
MLX 8-bit oraculumai/Manchego-MLX-8bit 4.5 GB 358b025b04001e50a065f8c87929175211264bd6182af74b67bd6caa2f639657
MLX 4-bit oraculumai/Manchego-MLX-4bit 2.4 GB e1bc5538b8dced2a857b4980dba045c2ca01db1c369aa416fc19f0f5c593e782

Each conversion is its own numeric series; the MLX cards carry their own measurements. No GGUF build is published.

Also here: LICENSE (Apache-2.0), NOTICE (every attribution and declaration), NATURAL_TASKS_ATTRIBUTION.md (the 101 Natural Instructions tasks and their 42 upstream sources), manchego_config.json (the decision contract; manchego-serve reads its contract), temperature_map.json (the serving map), contract_semif.py (the SemIf prompt, 2 to 16 options), contract_v2.py (the state-first prompt, 17 to 255 options), merge_record.json and eval/ (the aggregates behind this card).

Attribution and licences

Base model: Qwen3.5-4B (Apache-2.0, Alibaba Cloud). Training corpora, each used under its own licence: WANLI (Liu, Swayamdipta, Smith, Choi, 2022; CC BY 4.0); MultiNLI (Williams, Nangia, Bowman, 2018; mostly under the OANC licence; in the fiction genre one work is CC BY-SA 3.0, two are CC BY 3.0, the rest US public domain); Bitext customer support dataset (Bitext Innovations; CDLA-Sharing-1.0); GSM8K (Cobbe et al., 2021, OpenAI; MIT); MathDial (Macina et al., 2023, ETH Zurich; CC BY-SA 4.0); Spider (Yu et al., 2018, Yale LILY; CC BY-SA 4.0); ToolACE (Liu et al., 2024, Team-ACE; Apache-2.0); When2Call (Ross et al., 2025, NVIDIA; CC BY 4.0); NVD/CVE/CWE (NIST's National Vulnerability Database; CVE and CWE content copyright The MITRE Corporation, used under MITRE's terms); ProsocialDialog (Kim et al., 2022, Allen Institute for AI; CC BY 4.0); deepset prompt-injections (deepset; Apache-2.0); Super-NaturalInstructions (Wang, Mishra, et al., 2022; Apache-2.0 collection; each task's instances under its upstream dataset's licence, all listed in NATURAL_TASKS_ATTRIBUTION.md, where the one Gigaword-based task is noted). No row of any corpus is redistributed here. Doom states were recorded from ViZDoom (MIT) scenarios with Freedoom assets (BSD-3). The prompt format and system message are SemIf's direct-options prompt (SemIf, formerly OpenJev, by TheoLeeCJ; MIT). Evaluation sets: JevBench, the eight-dataset suite of logan-markewich/jeff (MIT) we call Public8, and Super-NaturalInstructions, each dataset under its own licence. Full notices: NOTICE. Illustration: the Manchego mascot, an AI-generated image (ChatGPT image generation) supplied by the project's author. The interface follows TypeSafe AI's System One contract; Manchego is an independent project, not affiliated with or endorsed by TypeSafe AI.

Downloads last month
91
Safetensors
Model size
5B params
Tensor type
BF16
·
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for oraculumai/Manchego

Finetuned
Qwen/Qwen3.5-4B
Finetuned
(837)
this model
Quantizations
2 models

Datasets used to train oraculumai/Manchego

Space using oraculumai/Manchego 1