Instructions to use oraculumai/Manchego with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use oraculumai/Manchego with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="oraculumai/Manchego")# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("oraculumai/Manchego") model = AutoModelForMultimodalLM.from_pretrained("oraculumai/Manchego", device_map="auto") - Notebooks
- Google Colab
- Kaggle

Manchego
A 4B model for typed decisions. One forward pass, a probability for every option.
Give Manchego a state (text or JSON), a question and a closed set of options. It returns a probability for each one:
pick an option (choice), a yes/no condition (noul), or an ordered level (score). It does not generate text.
This is Manchego v3: Qwen3.5-4B with a LoRA adapter (merged here), trained from the untrained base under the SemIf
prompt on about 100M tokens of decision rows. Manchego v2.1, the previous release, stays available at tag v2.1.
What changed in v3
- Trained from the base, not continued from v2.1. A fresh LoRA (rank 16) on the attention and gated-delta-net projections of every layer.
- A different prompt. SemIf's direct-options prompt: a system message and one JSON user message holding the evidence,
the criterion and lettered options. Serve v3 with it (
"contract": "semif", manchego-serve 0.2); v2.1's short prompt is not v3's. - More data. Hard computed families, answer judging, eight more public corpora (math, SQL, tool calling, security, dialogue safety, prompt injection) and a self-distilled replay toward the base.
- A serving temperature map that sharpens yes/no answers on purpose (below).
Results
| Suite (our runs, not official JevBench scores) | rows | untrained Qwen3.5-4B, v2.1's prompt | untrained Qwen3.5-4B, SemIf prompt | Manchego v2.1 | Manchego v3 | v3 minus v2.1 [95% interval] |
|---|---|---|---|---|---|---|
| JevBench, 231 public decisions (our run, not an official JevBench score) | 231 | 0.775 (0.493) | 0.801 (0.485) | 0.805 (0.423) | 0.857 (0.326) | +0.052 [+0.004, +0.102] |
| Public8, eight public real-text datasets | 1,600 | 0.686 (0.843) | 0.725 (0.847) | 0.732 (0.684) | 0.753 (0.609) | +0.021 [+0.002, +0.039] |
| 27 held-out Natural Instructions tasks (source-clean) | 1,561 | 0.678 (0.919) | 0.678 (1.056) | 0.701 (0.929) | 0.709 (0.941) | +0.007 [-0.016, +0.031] |
Accuracy with cross-entropy in brackets, temperature 1, one runtime for every model (see How the comparison was run). The JevBench rows on this card are our runs on its 231 public decisions, not official JevBench scores. JevBench's official score also reads sealed items, calibration, speed and cost, and only its maintainer runs it. v3 has no official JevBench result yet.
- JevBench's public decisions (our run, not an official JevBench score): v3 is ahead of v2.1 (+0.052, interval excludes zero; cross-entropy -0.097 [-0.171, -0.032]) and of the untrained base under v2.1's prompt (+0.082 [+0.022, +0.143]). The untrained base itself reads 0.801 under the SemIf prompt, so part of that step is the prompt.
- Public8: ahead of v2.1 (+0.021, resolved) with better probabilities (cross-entropy -0.075 [-0.107, -0.044]).
- Held-out task types: level with v2.1 (+0.007, interval through zero) and +0.031 [+0.002, +0.065] above the untrained base. Its probabilities there are no better than the base's (cross-entropy 0.941 against 0.919) and are overconfident (see Calibration).
Sealed sets, read once. Two sets held back for this read, with no recorded Manchego training or scoring on either, were opened once, on 2026-09-30, after v3 was fixed; nothing was selected by that read. The unseen-task set's record also notes earlier exposure: before it was reserved, Jev (TypeSafe AI's hosted decision service) was run on training rows of its tasks, and researchers saw some of its task names.
| Sealed set, read once | rows | untrained Qwen3.5-4B, v2.1's prompt | untrained Qwen3.5-4B, SemIf prompt | Manchego v2.1 | Manchego v3 | v3 minus v2.1 [95% interval] |
|---|---|---|---|---|---|---|
| 24 unseen Natural Instructions tasks (13 source clusters) | 5,397 | 0.641 (0.918) | 0.674 (0.911) | 0.659 (0.849) | 0.695 (0.861) | +0.042 [-0.018, +0.107] |
| Fresh out-of-distribution draws of four of the project's own computed families | 1,193 | 0.445 (1.161) | 0.508 (1.091) | 0.795 (0.466) | 0.762 (0.578) | -0.033 [-0.060, -0.005] |
The model columns pool rows. The unseen-task contrast is the mean over the 24 tasks of the per-task difference, with whole source clusters resampled; v3 is also +0.060 [-0.020, +0.143] above the untrained base there, again not resolved. The family draws (policy decisions, lookup chains, temporal reasoning, answer adequacy; 400 instances resampled) are where v2.1 kept training on v2's families and v3 started over from the base: v3 is 0.033 below v2.1 (resolved) and +0.317 [+0.280, +0.353] above the base.
The previous release's official JevBench result
| JevBench v1.5.4, official (maintainer-run) | score (95% interval) | rank | Intelligence | Calibration | Speed | Cost |
|---|---|---|---|---|---|---|
Manchego v2.1 (the previous release, tag v2.1) |
68.8 (59.6 to 70.3) | 11 of 106 | 51.2 | 84.9 | 88.6 | 64.3 |
Measured by JevBench's maintainer on their own offline GPU (RTX6000, bf16, temperature 1.0, original option order) with
manchego-serve v0.1.1 and oraculumai/Manchego at the v2.1 commit 77403228; the headline score weighs the four axes
equally. This row is v2.1's, not v3's. v3 has no official result yet.
Quick start
v3 is served by manchego-serve 0.2 (Apache-2.0). The server speaks
TypeSafe AI's System One wire contract (POST /v1/systemone), runs offline, and hashes the weights it loads and reports
whether they are the published ones. Pin this repository by its tag v3 (manchego-serve 0.2.0 pins the commit behind it).
git clone https://github.com/nschlaepfer/manchego-serve && cd manchego-serve && git checkout v0.2.0
# Docker, Linux + NVIDIA GPU: the build downloads the pinned weights once; the container runs offline
docker build -t manchego-serve:3-cuda --build-arg MODEL_REVISION=v3 .
docker run --rm --gpus all -p 127.0.0.1:8000:8000 manchego-serve:3-cuda \
--contract semif --temperature-map /models/manchego/temperature_map.json
# or pip (install the CUDA build of torch first on Linux + CUDA)
pip install .
manchego-serve-download --repo oraculumai/Manchego --revision v3 --out ./manchego-v3 # uses the network once
manchego-serve --model ./manchego-v3 --revision v3 --contract semif \
--temperature-map ./manchego-v3/temperature_map.json # offline from here on
curl -s localhost:8000/v1/systemone -H 'content-type: application/json' -d '{
"model": "manchego-3",
"state": "Customer: the blender I bought last week smells of burning and stopped working. Order 5521.",
"questions": {
"route": {"type": "choice", "instructions": "Which team should handle this message?",
"criteria": {"returns": "refunds, exchanges and defective items",
"shipping": "delivery status and lost parcels", "billing": null}},
"defective": {"type": "noul", "instructions": "Does the customer report a defective product?"},
"urgency": {"type": "score", "instructions": "How urgent is this message?",
"criteria": ["routine", "soon", "immediately"]}}}'
answers.route.probabilities holds one probability per option, answers.defective.noul is P(yes) and
answers.urgency.score is the expected level. Every response names the prompt each question got
(manchego.contract_by_question), the temperatures used and the weights' hash. manchego_config.json in this
repository already selects semif; --contract semif states it. Use --temperature-map off in place of the map
for probabilities at temperature 1, the setting of every number on this card.
Without the server (Transformers). Not a pipeline("text-classification") model: read the option-letter logits.
contract_semif.py in this repository renders the prompt v3 was trained on (2 to 16 options; the server handles 17 to
255 with the state-first prompt of contract_v2.py).
# pip install "transformers==5.17.0" torch huggingface_hub (tested: transformers 5.17.0, torch 2.10.0)
import importlib.util, torch
from huggingface_hub import hf_hub_download
from transformers import AutoModelForCausalLM, AutoTokenizer
REPO, REV = "oraculumai/Manchego", "v3"
spec = importlib.util.spec_from_file_location("contract_semif", hf_hub_download(REPO, "contract_semif.py", revision=REV))
semif = importlib.util.module_from_spec(spec); spec.loader.exec_module(semif)
tok = AutoTokenizer.from_pretrained(REPO, revision=REV)
model = AutoModelForCausalLM.from_pretrained(REPO, revision=REV, dtype=torch.bfloat16, device_map="auto").eval()
def decide(state, question, options, kind): # options: [(value, description or None)]; noul: [("true", None), ("false", None)]
r = semif.render(state, question, options, kind)
text = tok.apply_chat_template(r["messages"], tokenize=False, add_generation_prompt=True, enable_thinking=False)
ids = tok(text, return_tensors="pt", add_special_tokens=False).to(model.device)
with torch.no_grad():
logits = model(**ids).logits[0, -1].float()
z = logits[[tok.encode(L, add_special_tokens=False)[0] for L in r["letters"]]]
return dict(zip(r["keys"], torch.softmax(z, 0).tolist())) # divide z by the type's temperature to match the map
print(decide("WIN a free cruise! Reply YES now.", "Is this message spam?", [("true", None), ("false", None)], "noul"))
Apple silicon: Manchego-MLX-8bit and
Manchego-MLX-4bit.
Serving temperatures
| Question type | serving temperature (temperature_map.json) |
--temperature-map off |
|---|---|---|
| choice | 1.5 | 1.0 |
| noul | 0.2 | 1.0 |
| score | 1.0 | 1.0 |
The probabilities are softmax(z / T) over the offered option codes, with T set by the question type. A temperature never changes the chosen option.
The noul temperature 0.2 sharpens yes/no probabilities on purpose. JevBench v1.5 counts a yes/no answer whose P(yes) lies between 0.20 and 0.80 as wrong, so the map pushes answers out of that band. Under the map, P(yes) is therefore not a calibrated probability wherever the model is unsure. The choice temperature 1.5 flattens choice probabilities; score is unchanged.
Every number on this card is at temperature 1.
--temperature-map offserves exactly that.How it was chosen (TMAP-V15). Per question type, the temperature from a grid (0.2 to 2.0) that maximises an estimate of JevBench v1.5's composite, computed with v1.5's published scoring rules, on v3's records of development rows it never trained on. No JevBench item chose it:
TMAP-V15 fit choice noul score total development rows 10,485 16,050 1,647 28,182 The rule was registered before the fit, but after we had seen a noul-temperature sweep on those records and a report on JevBench's public items. After the fit, its effect on the public items was reported and changed nothing.
Which weights. The map is bound by hash to the three v3 builds: this repository's bf16 weights (
weights_sha256in Files), where it was fitted, and the MLX 8-bit and 4-bit builds, where it is applied as-is. manchego-serve 0.2 ships it as the default map for those hashes only; any other weights get temperature 1. It was fitted on records scored with the adapter applied to the base in padded batches, not on these merged weights' one-prompt logits; the merge moves option logits by at most 0.125 on its check rows.
Speed
v3's manchego_config.json turns on one piece of manchego-serve 0.2's CUDA fast path by default: the fast host path,
which is bit-for-bit the reference arithmetic with less host work (torch backend on CUDA only; --no-fast-host turns it
off). CUDA graphs stay off by default. On Linux, which is what the Docker recipe runs, padding each prompt to a graph
size made every prompt length slower:
| median latency, one decision per request (A10, Linux, Docker image, these weights, SemIf prompt) | reference | CUDA graphs + fast host |
|---|---|---|
| up to 256 prompt tokens | 54 ms | 71 ms |
| up to 512 | 134 ms | 143 ms |
| up to 1,024 | 190 ms | 275 ms |
| up to 2,048 | 357 ms | 543 ms |
On Windows (an RTX 5090) the same graphs were about five times faster than the reference (median 26 vs 124 ms), because
per-call overhead dominates there; --cuda-graphs turns them on if that is your platform. They pad each prompt to 256,
512, 1,024 or 2,048 tokens and replay a captured forward: on 83 invented test questions no chosen option changed, but
probabilities moved by up to 0.0385 (0.0626 across other bucket sets). Details: manchego-serve's docs/FAST_PATH.md. The
official Speed of 88.6 above is v2.1's, measured on the reference path.
Intended use and limits
- For: typed decisions over options your software supplies (routing, policy checks, triage, graded judgments, judging whether an answer is right), where you read the probabilities.
- Not for: unreviewed high-stakes decisions (medical, legal, financial, safety) without a person in the loop; open-ended text generation; images (inherited from the base and not measured).
- Options. v3 trained on 2 to 16 options under SemIf. Questions with 17 to 255 options, or with an empty state, go to
the state-first prompt of
contract_v2.py, which v3 never trained on. On the 104 rows of the held-out task set with more than 16 options, v3 and v2.1 each got 78 right. - Length. v3's longest training prompt was 6,386 tokens. Longer inputs are untested (manchego-serve accepts up to 32,768).
- Language: English. State: SemIf places the state in a JSON string field; v3's robustness to adversarial state was not measured.
- Machine-readable:
manchego_config.json.
Read before relying on it
- JevBench-informed. v2's computed families, which v3 retrains on, were designed from an earlier version's per-family JevBench hard-tier scores on the public items. v3's hard families were designed from JevBench's published hard-tier specification and family list, two of them also from this line's per-family public hard-tier results. No JevBench item text was opened, trained on or used to select rows.
- Public aggregate results. Development used JevBench's published aggregate results, never its items or per-item records: they ordered which new domain generators and public corpora were built first (they set no dose, label or selection), and v1.5's published scoring rules, with v2.1's official row, set the serving temperatures' objective (above).
- Phrase audit (inherited from v2). An early audit used Jev (TypeSafe AI's hosted decision service) to flag phrases in two template pools of v2's families; the flagged phrases were dropped. Jev never produced a label or a selection for v3.
- Labels. Targets are computed by our code, or a public dataset's original annotation, or (the self-distilled replay, 15% of the tokens) the untrained base model's own probabilities. ToolACE's and When2Call's labels were generated by their authors' model pipelines. No other label came from a language model, and no hosted decision service or teacher model of ours produced any.
- Text written by other companies' models. ProsocialDialog (utterances written by GPT-3), ToolACE (dialogues from unnamed generator models), When2Call (NVIDIA's generation pipeline) and MathDial (student turns written by gpt-3.5-turbo) contain model-written text. Their licences allow training and release with attribution; the account holder decided to release v3 with all four.
- MathDial's reproduction probe. On MathDial, v3 fails the project's trained-text reproduction rule (details under Provenance). The effect is larger on dialogues it never trained on, which points to the dialogues' style rather than memorised rows; the account holder accepted it for this release.
- Selection. v3 is the final update of its only run; nothing was selected inside it. Earlier candidates for v3 (among them a 300M-token run and an average of two adapters) did not pass their registered development reads. The sealed and public reads above came after, once, and selected nothing.
Training
| Base | Qwen/Qwen3.5-4B @ 851bf6e8 (hybrid gated-delta-net + attention), untrained |
| Method | LoRA, rank 16, alpha 32, on q/k/v/o of the 8 attention layers and the five in/out projections of the 24 gated-delta-net layers (152 modules, 14,376,960 trainable parameters), merged into the base for release |
| Objective | cross-entropy on the offered option letters' logits at the last prompt position, with exact soft targets where a row declares a distribution; no loss on any text token |
| Prompt | SemIf's direct-options prompt on every row (2 to 16 options); choice options shuffled at every presentation, noul true then false |
| Run | 255,991 rows in one pass, 8,771 updates (token-budget batches of 12,288 padded tokens, at most 64 rows), 99,362,993 prompt tokens; lr 3e-5, 85 warm-up updates, cosine to 0, AdamW, no weight decay, gradient clip 1; one NVIDIA GH200, 2026-09-28, 00:56 to 04:48 UTC; every planned row consumed |
| Checkpoint | the final update; nothing was selected inside the run |
| Training data (bucket) | share of the estimated prompt tokens | rows | labels |
|---|---|---|---|
| Computed decision families of v2 (policies, lookups, temporal, adequacy, routing, skills, arena and Doom states, score levels), regenerated | 27.3% | 48,914 | computed by our generators |
| Hard computed families (long policy documents, multi-hop lookups, abstention, temporal strata, exact probabilities, paraphrase invariance, policy compliance) | 22.7% | 19,214 | computed |
| Answer judging (judging worked answers; judge-error repair) | 8.0% | 22,547 | computed |
| Human-annotated text: WANLI, MultiNLI, Bitext; 101 Natural Instructions tasks; GSM8K, MathDial, Spider, ToolACE, When2Call, NVD/CVE/CWE, ProsocialDialog, deepset prompt-injections | 20.0% | 83,701 | the datasets' original annotation (ToolACE and When2Call: their authors' generated labels) |
| Self-distilled replay (a partition of the same public corpora, and fresh answer-judging instances) | 15.0% | 61,739 | the untrained Qwen3.5-4B's own probabilities over the options |
| Domain generators (auth-log triage, code output, exact draws, genetic crosses, limitation deadlines, loan amortisation, math-answer grading, plan cost sharing) | 5.0% | 11,087 | computed |
| Calibration rows | 2.0% | 8,789 | computed |
| Total | 100% | 255,991 |
By target, 178,334 rows train toward one gold option, 15,918 toward an exact distribution and 61,739 toward the base's own probabilities. By type: 158,909 choice, 84,115 noul and 12,967 score rows.
No row comes from a benchmark on this card. Every source went through an 8-word overlap gate against the evaluation sets (JevBench's public items, Public8, the held-out and sealed task sets) before the training file was built, and the rows that matched were dropped. Public8's eight datasets and the held-out tasks' upstream datasets appear nowhere in the declared training data (checked by name); what the base model saw in pretraining is unknown. GSM8K and When2Call are also Decision Index benchmarks: v3 trained on their train splits only, and their test splits were kept out of training. "Computed" means exact under the task definition as implemented.
Calibration
| Calibration error (ECE, temperature 1) | untrained Qwen3.5-4B, v2.1's prompt | untrained Qwen3.5-4B, SemIf prompt | Manchego v2.1 | Manchego v3 | v3 mean confidence | v3 accuracy |
|---|---|---|---|---|---|---|
| JevBench, 231 public decisions (our run, not an official JevBench score) | 0.062 | 0.059 | 0.072 | 0.055 | 0.854 | 0.857 |
| Public8 | 0.122 | 0.122 | 0.053 | 0.055 | 0.809 | 0.753 |
| 27 held-out Natural Instructions tasks | 0.107 | 0.119 | 0.083 | 0.129 | 0.829 | 0.709 |
| 24 unseen Natural Instructions tasks (sealed) | 0.151 | 0.148 | 0.098 | 0.141 | 0.836 | 0.695 |
| Fresh out-of-distribution draws of the project's families (sealed) | 0.141 | 0.075 | 0.051 | 0.050 | 0.800 | 0.762 |
On familiar ground v3 is as well calibrated as v2.1. On unfamiliar task definitions it is overconfident: on the 27 held-out tasks its mean confidence is 0.829 against an accuracy of 0.709, and its calibration error (0.129) is worse than v2.1's (0.083) and the base's. v2.1's second stage had repaired exactly this in v2; v3, trained from the base, has it again. The serving map's choice temperature of 1.5 flattens choice probabilities; its effect on these sets was not measured.
Weaknesses
- Not more accurate than v2.1 on unfamiliar task types (+0.007 on 27 held-out tasks, +0.042 on 24 sealed unseen tasks, both through zero), and overconfident there (above).
- Below v2.1 on the project's own families out of distribution (-0.033, resolved, on sealed fresh draws), and on 200 human-labelled rows from the development splits of v2.1's real-text corpora (0.785 against v2.1's 0.910, MLX 8-bit; see the MLX cards). A quarter of those rows come from banking77, which v3 did not train on; v2.1 fitted v2's families and those corpora more closely.
- Large menus go through a prompt v3 never trained on (17 to 255 options, or an empty state).
- Yes/no probabilities under the serving map are sharp by design, not calibrated. Use
--temperature-map offwhen you need P(yes) as a probability. - JevBench's public items may not predict its sealed items. The step on the 231 public decisions is our run; the official score also reads sealed items, and the project's own instruments have not predicted JevBench's Intelligence axis.
- One run, one seed. No second seed of v3 was trained.
- Date and number arithmetic in one pass is unreliable outside its own templates. Do the arithmetic in code and ask the model the judgment.
- Numerics of serving. Scoring several prompts in one padded batch moves probabilities slightly at bf16 or 8 bits; score one prompt at a time (manchego-serve's default) when exact reproducibility matters.
- A 4B model. Its world knowledge is the base model's; it will be wrong on questions that need facts it does not have.
How the comparison was run
Every model in the tables above ran in one runtime on one RTX 5090: bf16, the adapters applied to the base (not merged),
padded batches of up to 8,192 tokens, identical rows, temperature 1. v3 ran under the SemIf prompt (the state-first
prompt beyond 16 options, as manchego-serve 0.2 serves it); v2.1 and the untrained base ran under v2.1's served policy
(the short prompt up to 26 options, state-first beyond), and the base also under SemIf. v2.1 and the untrained base
reproduced their archived records on the three public suites (the runtime check passed). Intervals are paired 95%
bootstraps: JevBench resamples its scenario groups, Public8 its items within each dataset, the held-out set whole tasks.
The merged weights in this repository reproduce base plus adapter within the merge check (merge_record.json: at most
0.125 in any option logit on 12 rows, no changed decision); a bf16 merge is not bit-identical. Aggregates:
eval/.
Provenance
- Adapter: update 8,771, the final update of a fresh LoRA trained from the untrained Qwen3.5-4B;
adapter_model.safetensorssha256c43e2688…(full hash inmerge_record.json). Training file sha256240b5cbc…, 255,991 rows, checked row by row on the training machine by a launch gate: no reserved evaluation task; no target from an external teacher model or a hosted decision service (the replay's targets are the base model's own); every row's permission recorded. - A release check walks the permission record of every training row. It cleared v3 on 2026-09-30, after the account
holder's decisions above;
NOTICEcarries every attribution those permissions require. - Reproduction probe (verbatim continuation of training texts under teacher forcing, the adapter against the base,
texts trained on against held-out texts;
eval/de_minimis_probe.json): Bitext, MultiNLI fiction and Spider pass every rule. MathDial fails two: under the decision prompt, 0.075 of trained dialogues are continued verbatim for 8 tokens or more against the base's 0.045 (the rule allows 0.02 more), and 6 trained dialogues are continued for 12 tokens or more by the adapter alone (the rule allows none; as raw text, 2). On 200 MathDial dialogues it never trained on, the same measures read 0.155 against 0.085, and 14. - Weights:
model.safetensors-*in this repository are the base checkpoint with every language-model tensor replaced by its merged value; the vision tower, the multi-token-prediction head, the config and the tokenizer are the base model's. Manchego was trained and evaluated on text only.
Versions
This repository's main holds Manchego v3 (2026-09-30), tagged v3. Manchego v2.1 (2026-09-21) stays available,
unchanged, at tag v2.1 (revision="v2.1"): its weights, card, DETAILS.md and evidence. v2.1 needs its own prompts
(manchego-serve's contract auto); v3 needs SemIf. Earlier versions are not published.
Files
| Format | repository | size | weights_sha256 (as manchego-serve reports it) |
|---|---|---|---|
| bf16 (Transformers) | oraculumai/Manchego |
9.3 GB | 2ee838433bfe278a226dc644667ad4a99ece82cc47325c7645a7dae723c1863b |
| MLX 8-bit | oraculumai/Manchego-MLX-8bit |
4.5 GB | 358b025b04001e50a065f8c87929175211264bd6182af74b67bd6caa2f639657 |
| MLX 4-bit | oraculumai/Manchego-MLX-4bit |
2.4 GB | e1bc5538b8dced2a857b4980dba045c2ca01db1c369aa416fc19f0f5c593e782 |
Each conversion is its own numeric series; the MLX cards carry their own measurements. No GGUF build is published.
Also here: LICENSE (Apache-2.0), NOTICE (every attribution and declaration), NATURAL_TASKS_ATTRIBUTION.md (the
101 Natural Instructions tasks and their 42 upstream sources), manchego_config.json (the decision contract; manchego-serve
reads its contract), temperature_map.json (the serving map), contract_semif.py (the SemIf prompt, 2 to 16 options),
contract_v2.py (the state-first prompt, 17 to 255 options), merge_record.json and eval/ (the aggregates behind this
card).
Attribution and licences
Base model: Qwen3.5-4B (Apache-2.0, Alibaba Cloud). Training corpora, each used under its own licence: WANLI (Liu,
Swayamdipta, Smith, Choi, 2022; CC BY 4.0); MultiNLI (Williams, Nangia, Bowman, 2018; mostly under the OANC licence;
in the fiction genre one work is CC BY-SA 3.0, two are CC BY 3.0, the rest US public domain); Bitext customer support
dataset (Bitext Innovations; CDLA-Sharing-1.0); GSM8K (Cobbe et al., 2021, OpenAI; MIT); MathDial (Macina et
al., 2023, ETH Zurich; CC BY-SA 4.0); Spider (Yu et al., 2018, Yale LILY; CC BY-SA 4.0); ToolACE (Liu et al.,
2024, Team-ACE; Apache-2.0); When2Call (Ross et al., 2025, NVIDIA; CC BY 4.0); NVD/CVE/CWE (NIST's National
Vulnerability Database; CVE and CWE content copyright The MITRE Corporation, used under MITRE's terms); ProsocialDialog
(Kim et al., 2022, Allen Institute for AI; CC BY 4.0); deepset prompt-injections (deepset; Apache-2.0);
Super-NaturalInstructions (Wang, Mishra, et al., 2022; Apache-2.0 collection; each task's instances under its upstream
dataset's licence, all listed in NATURAL_TASKS_ATTRIBUTION.md, where the one Gigaword-based task is noted). No row of
any corpus is redistributed here. Doom states were recorded from ViZDoom (MIT) scenarios with Freedoom assets (BSD-3).
The prompt format and system message are SemIf's direct-options prompt (SemIf, formerly OpenJev, by TheoLeeCJ; MIT).
Evaluation sets: JevBench, the eight-dataset suite of logan-markewich/jeff (MIT) we call Public8, and
Super-NaturalInstructions, each dataset under its own licence. Full notices: NOTICE. Illustration: the Manchego mascot,
an AI-generated image (ChatGPT image generation) supplied by the project's author. The interface follows TypeSafe AI's
System One contract; Manchego is an independent project, not affiliated with or endorsed by TypeSafe AI.
- Downloads last month
- 91