snap1-2b

GitHub · Website · Ollaya

A 2B model that answers from its logits. snap1 is MiniCPM5-2B fine-tuned for one job: a typed decision in a single forward pass. It is the model for snap, the engine that turns a language model into a decision function. snap asks a question, reads one row of logits at the answer position, and returns a probability for every option. The model never generates text.

Parameters 2B (MiniCPM5-2B base)
Files snap1-2b-q4_k_m.gguf (1.5 GB, default), snap1-2b-q8_0.gguf (2.7 GB), snap1-2b-bf16.gguf (5.0 GB)
Answers choice, yes/no, score levels, numeric ranges, optional "cannot tell"
Languages English, Italian
License Apache-2.0

Quickstart

brew install emnlmn/snap/snap          # or a release tarball: github.com/emnlmn/snap/releases
snap serve --model snap1-2b            # downloads once, then runs offline
curl -s localhost:8018/v1/systemone -d '{
  "state": "Hi, I was charged twice for March. Please refund the duplicate today.",
  "questions": {
    "route":  {"type": "choice", "criteria": {"billing": "payments, refunds", "tech": "bugs, outages"}},
    "urgent": {"type": "noul", "instructions": "Does this need a same-day reply?"}
  }
}'

The API is POST /v1/systemone, wire-compatible with TypeSafe Jev. Playground: http://localhost:8018/playground.

Also on Ollaya: ollaya pull snap:2b (Q8_0, needs Ollaya 0.10.0 or newer). Same prompt, checked byte for byte against snap; typed-decisions accuracy there is 0.648 against 0.655 in snap, because Ollaya reads the bare option letter and snap takes the highest logit among the letter's bare, space- and newline-prefixed tokens.

This checkpoint is trained on snap's exact prompt bytes. Use it through snap or Ollaya. In a plain chat runtime it is just a slightly odd chat model.

Results

All numbers were measured with snap 0.5.0, Q4_K_M, and compared paired against the base model on the same cases.

MiniCPM5-2B (base) snap1-2b
typed-decisions test, accuracy, zero-shot 0.486 0.655
typed-decisions, KL from gold ↓ 1.673 0.292
typed-decisions, Brier ↓ 0.464 0.159
typed-decisions, calibration error (ECE) ↓ 0.316 0.041
snap eval suite (304 scored, never trained on) 0.697 0.852
typed-decisions, 5-question request, median, Apple M1 Max 607 ms 534 ms
typed-decisions, 5-question request, median, RTX 4090 — 48 ms

Zero-shot: this checkpoint never saw the typed-decisions data, neither the test split nor the train split, nor its four workflows.

The typed-decisions probabilities are raw, with no calibration applied. The fine-tune changes weights, not compute, so it runs at the base model's speed; the two M1 Max timings come from the same session and differ by run-to-run noise. The RTX 4090 row is the median on a fast CUDA host (the base model was not re-run there); latency depends on the host. The same 0.5.0 workload on CUDA scored 0.660 accuracy; the CUDA/Metal difference is run noise.

The fine-tune also makes answers more stable. On snap's eval suite, with the options reversed, irrelevant text injected or the question rephrased, the answer holds 86% of the time, against 76% on the base model. It also says "cannot tell" exactly when it should: on the suite's abstain cases it catches 16 out of 16 with no false abstention (base model: 13 out of 16).

Part of the wider gap is the renderer, not the weights. Since 0.4.0, snap renders structured state as TOON (the YAML-lite variant we benchmarked answered identically on every case). On this rendering the base drops (0.502 → 0.486 on typed-decisions) while snap1, trained on those exact prompt bytes, gains (0.624 → 0.655). In 0.5.0 the prompt also renders yes/no criteria as explicit outcome lines, which is where most of the latest gain comes from (+4 points on noul questions).

Limits

  • Numeric questions are the weakest type, especially for values at the edge of the range or outside it. Check them on your own data.
  • A 2B model reads short states well and long documents less well. For multi-page inputs, measure before you rely on it.
  • Comparing numbers across fields is the weakest skill. On typed-decisions' invoice and agent-trace workflows, which match amounts and events against a document, it scores 0.622, against 0.688 on the other two workflows.
  • It does not compute. Questions that need arithmetic, such as the days between two dates, are unreliable. Its probabilities usually show it: they come out flat instead of confident.
  • Calibration depends on the domain. On typed-decisions the raw probabilities are already well calibrated, but on your own data they can drift. snap calibrate fits per-type temperatures on your own labeled cases.
  • Don't use it as the only check on high-stakes decisions (medical, legal, credit) without human review.

Training

snap1 is a LoRA (rank 16) on MiniCPM5-2B, merged and converted to GGUF. The loss is a softmax over the legal option letters only, read the same way snap reads them. The data mixes public decision datasets with in-house synthetic cases, and every row goes through schema, dedup and contamination checks against snap's eval suite. Rows whose license does not allow commercial use were dropped.

Two parts of the data fight shortcut learning. The corpus includes 1,100 contrastive families (4,510 rows): near-identical states where a single deciding field flips — a threshold crossing, a flag — so the correct answer changes and the label has to come from reading that field. And at export, every choice question's options are permuted; the order the data was generated in is never seen, so the option order carries no signal about the answer.

Public sources:

Dataset Licenses of the rows used
ZefanCai/Open-Jev CC0-1.0
tasksource/tasksource-jev-typed-decisions per row: mostly Apache-2.0, MIT, CC-BY and CC0; some CC-BY-SA and GPL rows, and rows with no license stated; non-commercial rows dropped
Praveenrajus/jev-bench per config: CC-BY-4.0 (incl. HelpSteer2), MIT, CC0, CC-BY-SA (e.g. ARC, STS benchmark), Google PAWS, and configs with no license stated (SST, UCI SMS Spam); research-only configs dropped
carlomarxx/trilemma-of-truth CC-BY-4.0

Of the public rows, about 72% carry permissive licenses, 10% share-alike or copyleft licenses (CC-BY-SA, GPL) and 18% no license stated in their source dataset. The public rows are about half of the training data; the rest is in-house synthetic.

The training pipeline is open source: training/, documented in TRAINING.md.

Citation

@misc{snap1-2026,
  title  = {snap1: a 2B model that answers from its logits},
  author = {Logitlab},
  year   = {2026},
  url    = {https://huggingface.co/logitlab/snap1-2b-GGUF}
}
Downloads last month
967
GGUF
Model size
3B params
Architecture
llama
Hardware compatibility
Log In to add your hardware

4-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for logitlab/snap1-2b-GGUF

Quantized
(102)
this model
Finetunes
1 model

Datasets used to train logitlab/snap1-2b-GGUF