Leo 1.7B (leo-1.7b-v3)

Leo is an open-weight decision model. You send a state (text or JSON) and typed questions (choice, score, or noul yes/no). It returns a calibrated probability distribution for every question from a single forward pass, with no text generation. Typical uses: routing, classification, moderation triage, grading on a rubric, policy checks, and the next-action choices of an agent.

The request and response format follows the POST /v1/systemone API that TypeSafe documents for its Jev models, so existing clients work after a base-URL change. Leo is an independent project: it is not affiliated with TypeSafe, contains none of its weights, and was not trained on Jev outputs. Jev was called only to score it next to Leo on identical requests.

  • Code, training pipeline, benchmarks: github.com/SuparvaCode/leo
  • Versions in this repo: main = leo-1.7b-v3 (recommended), v4-experimental = leo-1.7b-v4

Model details

Developer Suparva Baranwal
Version leo-1.7b-v3
Model type Decoder-only transformer run prefill-only, LoRA adapter, listwise pointer head
Base model Qwen/Qwen3-1.7B-Base (Apache-2.0), revision ea980cb0a6
Parameters about 1.7B in the base; 20.3M trained (LoRA r16 on every attention and MLP projection, plus markers and head)
Files adapter/ (LoRA, 70 MB), leo_head.safetensors (head and marker embeddings, 12 MB), leo_config.json (calibration and metadata), leo/ (inference code)
Inputs a state (string, JSON object or array) and many typed questions per request
Outputs per question: choice + probabilities + confidence, or score + probabilities + legend, or noul = P(yes)
Context trained on states up to 2,048 tokens; browser states up to about 12k tokens were used in evaluation
Languages English plus 50+ languages via MASSIVE, SIB-200 and multilingual sentiment data; much weaker on low-resource languages
Licence Apache-2.0 (weights and code); training datasets keep their own terms

Results at a glance

Measured on one machine, with Jev 1.13 (live API) (jev-1.13.0, September 2026) answering exactly the same requests. Per-task tables and the protocol follow further down.

benchmark leo-1.7b-v3 leo-1.7b-v4 Jev 1.13 (live API)
Browser tasks passed (jev-ultrafast, 7 tasks x 3 runs) 18/21 12/21 18/21 (one per matched session)
False DONE rate, held-out browser screens (lower is better) 0.067 0.000 not measured
JevBench, all 231 items 0.697 0.680 0.861
JevBench, standard tier 0.958 0.903 0.986
JevBench, hard tier 0.396 0.396 0.721
Held-out classification, mean accuracy (4 datasets, zero-shot) 0.628 0.627 0.689
Multilingual, 2,049 items in 15 languages (blind) 0.560 0.557 0.857
Calibration error (ECE), held-out mean (lower is better) 0.099 0.094 0.167

In short: leo-1.7b-v3 ties Jev on the browser suite (18/21 vs 18/21), is close on JevBench's easy and standard tiers, and beats Jev on tweet topics. It is clearly behind on hard reasoning (long policies, ambiguity, trade-offs, multi-hop), knowledge-heavy exams and low-resource languages. Most of that gap comes from the size of the 1.7B base model.

The v4 experiment: fixing false DONE

Problem in v3. leo-1.7b-v3 declares a browser task DONE once its action history covers every part of the goal, whatever the page shows. On Google Flights it stopped with DONE in all three runs without completing the search, with P(DONE) = 0.986 while only a seat-class list was on screen, and 0.999 on a blank render.

What v4 changed. leo-1.7b-v4 continues v3's weights (1 epoch, half the learning rate) on 30,000 replayed v3 requests plus 9,000 steps from a new simulator family (leo/data/browser_evidence.py) in which DONE is only right with visible evidence: close the open dropdown, wait out blank or loading renders, resubmit when results show an older search, report BLOCKED when a requested value is not offered.

What happened.

  • The targeted flaw is fixed in simulation and largely on the real site. The false-DONE rate on held-out screens fell from 0.067 to 0.000. On live Google Flights, v4 no longer claims success: all 3 runs ended with an honest BLOCKED (the seat-class state that fooled v3 went from P(DONE) 0.986 to 0.655).
  • It introduced a new failure. On jev-ultrafast's two reading-room tasks (open one article, no search) v4 opens the right article and then goes back to the list, over and over: 0/6 passed, every run hit the 60-action budget. The likely cause: every goal in the new family includes a search, so an item page reached without a search in the history always meant "go back" in training. Browser total: 18/21 for v3, 12/21 for v4.
  • JevBench fell from 0.697 to 0.680 (standard tier 0.958 to 0.903), mostly policy and trap items where v4 now answers "yes" or skips "other"/"unknown". Held-out classification and multilingual accuracy are unchanged within noise.

Decision. leo-1.7b-v3 stays the recommended model on main. leo-1.7b-v4 is published on the v4-experimental branch for research and for agent loops that need an honest BLOCKED more than they need open-an-item tasks. The fix for the next version is to mix goals without a search step into the evidence family and to re-check the reading-room tasks before release.

Quick start

pip install torch transformers peft safetensors numpy pydantic huggingface_hub
import sys
from huggingface_hub import snapshot_download

path = snapshot_download("Suparva/leo-1.7b")   # revision="v4-experimental" for v4
sys.path.insert(0, path)            # the repo ships its own inference code in leo/
from leo.infer import Leo

leo = Leo.load(path, dtype="bf16")  # use "fp32" on CPU; the Qwen3 base is fetched on first use
out = leo.system_one(
    {
        "ticket": {
            "subject": "Charged twice",
            "body": "I was billed twice this month. Please fix it today."
        }
    },
    {
        "route": {
            "type": "choice",
            "instructions": "Which team should handle `ticket`?",
            "criteria": {
                "billing": "charges, refunds, invoices",
                "technical": "bugs, outages",
                "other": None
            }
        },
        "urgent": {
            "type": "noul",
            "instructions": "Does the customer need an answer today?"
        },
        "anger": {
            "type": "score",
            "instructions": "How upset is the customer?",
            "criteria": [
                "calm",
                "annoyed",
                "very angry"
            ]
        }
    },
)
print(out)

Real output of this request (the exported folder, loaded on its own). The Python API returns 4 decimals; the HTTP server returns TypeSafe's exact shape by default (2 decimals summing to exactly 1).

{
  "model": "leo-1.7b-v3",
  "answers": {
    "route": {
      "type": "choice",
      "choice": "billing",
      "probabilities": {
        "billing": 0.9026,
        "technical": 0.0374,
        "other": 0.06
      },
      "confidence": 0.8539
    },
    "urgent": {
      "type": "noul",
      "noul": 0.8721
    },
    "anger": {
      "type": "score",
      "score": 1.1116,
      "legend": {
        "0": "calm",
        "1": "annoyed",
        "2": "very angry"
      },
      "probabilities": {
        "0": 0.0432,
        "1": 0.802,
        "2": 0.1548
      },
      "confidence": 0.7029
    }
  },
  "usage": {
    "input_tokens": 90,
    "output_tokens": 136
  },
  "latency_ms": 847.38
}

leo.predict_many([{"state": ..., "questions": ...}, ...]) batches many requests.

HTTP server

pip install fastapi uvicorn httpx
cd <snapshot path>
python -m leo.serve --model . --port 8000                          # binds 127.0.0.1
LEO_API_KEY=change-me python -m leo.serve --model . --host 0.0.0.0 # a key is required off loopback

POST /v1/systemone and GET /v1/models follow TypeSafe's API: bearer auth whenever LEO_API_KEY is set, 422 on invalid requests, body-size and question-count limits. The server refuses to bind a public address without a key. Flags:

  • --order-views 2 averages each choice over two option orders: about half the option-order sensitivity, for 1.4 to 2.8 times the latency.
  • --dtype fp32 gives answers that do not depend on which other questions share the request (bf16 moves probabilities by up to about 0.02).
  • --precise returns 4-decimal probabilities plus latency_ms instead of the TypeSafe-exact shape.

leo.client.SystemOneClient talks to a Leo server or any other /v1/systemone endpoint.

Question types

type criteria answer
choice object mapping option key to a description (string, object, array or null), 2 to 255 options choice, probabilities over the keys, confidence = (K·p_max − 1)/(K − 1)
score ordered array of 2 to 10 level descriptions score = Σ i·pᵢ, probabilities, confidence, legend
noul optional {"true": ..., "false": ...} noul = P(yes)

Instructions can be a string or any JSON value and can refer to parts of the state by path, such as `ticket.body`. Question IDs are never shown to the model. Questions cannot see each other, so adding or removing one does not change the others.

How it works

  • Backbone: Qwen/Qwen3-1.7B-Base run prefill-only (no decoding), adapted with LoRA r16 / alpha 32.
  • Packed layout: the state is encoded once; each question follows it in the same sequence under a block-causal mask, with position ids restarting after the state. Questions share the state encoding but cannot attend to each other.
  • Reserved markers: state, question, option and decision boundaries are trainable embeddings outside the vocabulary, so text inside the state cannot forge them.
  • Listwise pointer readout: a small head scores each option's end marker against the question's decision marker, which comes after the full option list, so options are judged together.
  • Proper-scoring-rule training: log loss against the label distribution plus a ranked probability score term for score questions.
  • Calibration: one temperature per question type and option-count bucket, fitted on dev data (dev top-label ECE 0.014). Confidence uses the formulas TypeSafe publishes.

Evaluation

  • Every Jev number comes from live calls on byte-identical requests (same state, instructions, option keys and option order), September 2026. Jev responses were used only for scoring, never for training, prompt tuning or checkpoint selection. Checkpoints were selected on dev loss alone.
  • Held-out classification and the multilingual suite are blind: none of those datasets or task families were trained on. JevBench items were read while designing the data, so treat that score as seen.
  • The browser suite runs TypeSafe's own agent unchanged and checks the final page independently. Live sites (Wikipedia, Google Flights) can change between sessions; Jev's own Google Flights result differed between the two sessions.
  • TYPE_TEXT values in the browser suite come from the same local Qwen3-1.7B helper for both arms.
benchmark leo-1.7b-v3 leo-1.7b-v4 Jev 1.13 (live API)
Browser tasks passed (jev-ultrafast, 7 tasks x 3 runs) 18/21 12/21 18/21 (one per matched session)
False DONE rate, held-out browser screens (lower is better) 0.067 0.000 not measured
JevBench, all 231 items 0.697 0.680 0.861
JevBench, standard tier 0.958 0.903 0.986
JevBench, hard tier 0.396 0.396 0.721
Held-out classification, mean accuracy (4 datasets, zero-shot) 0.628 0.627 0.689
Multilingual, 2,049 items in 15 languages (blind) 0.560 0.557 0.857
Calibration error (ECE), held-out mean (lower is better) 0.099 0.094 0.167

Browser automation

browser-use/jev-ultrafast is TypeSafe's own open-source browser agent: each step sends one /v1/systemone request with an operation question and one target question per operation. The agent loop, request builder, response validation, DOM snapshot and executor were left unchanged; only the model answering the request differs. A run passes only when the agent stops with DONE and independent checks on the final page hold (a DONE alone is never trusted). 3 runs per task; each Leo version had its own session with Jev, the two arms alternating within each repeat (same Chrome profile, viewport, 60-action budget, 120 s limit). Five tasks use jev-ultrafast's local fixture, two use live websites.

Cells are passed runs, then median time / model decisions.

task leo-1.7b-v3 Jev (same session) leo-1.7b-v4 Jev (same session)
travel-casa-flora 3/3 · 5.0 s / 6 3/3 · 3.6 s / 6 3/3 · 5.3 s / 6 3/3 · 3.5 s / 6
research-finite-choices 3/3 · 5.7 s / 12 3/3 · 0.8 s / 2 0/3 · 33.0 s / 61 3/3 · 0.9 s / 2
travel-glasshouse 3/3 · 5.0 s / 6 3/3 · 3.5 s / 6 3/3 · 5.3 s / 6 2/3 · 3.5 s / 6
travel-serra-lodge 3/3 · 4.3 s / 5 3/3 · 3.0 s / 5 3/3 · 4.4 s / 5 3/3 · 3.1 s / 5
research-confidence 3/3 · 0.7 s / 2 3/3 · 0.8 s / 2 0/3 · 31.6 s / 61 3/3 · 0.9 s / 2
wikipedia-godel (live site) 3/3 · 11.6 s / 5 3/3 · 4.0 s / 4 3/3 · 10.1 s / 4 3/3 · 4.3 s / 5
google-flights (live site) 0/3 · 18.7 s / 14 0/3 · 1.6 s / 3 0/3 · 48.5 s / 36 1/3 · 13.5 s / 17
all 18/21 18/21 12/21 18/21

How the failed runs ended:

  • leo-1.7b-v3, google-flights: 3 run(s) stopped with DONE
  • Jev (session with leo-1.7b-v3), google-flights: 3 run(s) stopped with BLOCKED
  • leo-1.7b-v4, research-finite-choices: 3 run(s) hit the 60-action budget
  • leo-1.7b-v4, research-confidence: 3 run(s) hit the 60-action budget
  • leo-1.7b-v4, google-flights: 3 run(s) stopped with BLOCKED
  • Jev (session with leo-1.7b-v4), travel-glasshouse: 1 run(s) stopped with DONE
  • Jev (session with leo-1.7b-v4), google-flights: 2 run(s) stopped with BLOCKED

False-DONE probe

scripts/done_probe.py replays 400 simulated browser steps from two website themes that never appear in any training set, on screens where the action history can look finished while the page proves nothing.

screen n leo-1.7b-v3 false DONE leo-1.7b-v4 false DONE leo-1.7b-v3 step accuracy leo-1.7b-v4 step accuracy
blank page between two renders (right move: WAIT) 18 0.556 0.000 0.444 1.000
opened item (DONE is right if it matches) 18 n/a n/a 0.667 1.000
settled results page (DONE is right if it matches) 45 0.000 0.000 0.822 1.000
results still loading (WAIT) 18 0.111 0.000 0.889 1.000
a dropdown list covers the page 82 0.085 0.000 0.756 1.000
a counter pop-over with its own Done button 65 0.031 0.000 0.723 1.000
results still show the previous search 19 0.053 0.000 0.474 1.000
search form not submitted yet 135 0.015 0.000 0.556 1.000
all screens 400 0.067 0.000 0.665 1.000

Real Google Flights states from leo-1.7b-v3's recorded runs, replayed as sent:

recorded state leo-1.7b-v3 P(DONE) / answer leo-1.7b-v4 P(DONE) / answer
26 elements: Skip to main content / Accessibility feedback / Explore 0.378 / CLICK 0.001 / CLICK
4 elements: Economy / Premium economy / Business / First 0.986 / DONE 0.655 / DONE
blank page (0 elements) 0.999 / DONE 0.092 / WAIT

None of these pages show flight results, so DONE is wrong on all of them. The held-out themes come from the same simulator family as v4's new training data, so the probe is easier than real sites; the browser suite above is the real test.

JevBench (public items)

231 typed decisions from fstandhartinger/jevbench (commit 1bcc55eb6c). JevBench items were read while designing the training data, so this is not a blind score.

system easy (48) standard (72) hard (111) all (231) ECE
leo-1.7b-v3 1.000 0.958 0.396 0.697 0.101
leo-1.7b-v4 1.000 0.903 0.396 0.680 0.132
Jev 1.13 (live API) 1.000 0.986 0.721 0.861 0.057
Accuracy by task family
tier / family n leo-1.7b-v3 leo-1.7b-v4 Jev 1.13 (live API)
easy/extraction 12 1.00 1.00 1.00
easy/fact 12 1.00 1.00 1.00
easy/intent 12 1.00 1.00 1.00
easy/tool_selection 12 1.00 1.00 1.00
hard/adversarial 6 0.33 0.33 1.00
hard/ambiguous 7 0.14 0.14 0.86
hard/judge_hard 17 0.53 0.53 0.71
hard/long_policy 19 0.21 0.32 0.63
hard/multi_hop 18 0.44 0.50 0.83
hard/probability 10 0.40 0.30 0.70
hard/routing_hard 5 1.00 1.00 1.00
hard/temporal_numeric 15 0.20 0.20 0.27
hard/tradeoff 6 0.17 0.17 0.83
hard/trap 8 0.88 0.62 1.00
standard/adequacy 12 0.83 0.83 1.00
standard/extraction 12 1.00 0.92 1.00
standard/intent 12 1.00 0.92 1.00
standard/ordinal 12 1.00 1.00 1.00
standard/policy 12 1.00 0.83 0.92
standard/routing 12 0.92 0.92 1.00

Held-out classification (zero-shot)

Four datasets whose sources and task families were never trained on, with the protocol of elcronos/jev-vs-open-decision-models: emotion (6 labels), tweet_topic (6), fin_topic (20 financial-news topics), daily_dialog (7 dialogue emotions, 82% "no emotion"). Cells are accuracy / macro-F1 / ECE.

system emotion tweet_topic fin_topic daily_dialog mean accuracy
leo-1.7b-v3 0.564 / 0.469 / 0.059 0.845 / 0.703 / 0.123 0.456 / 0.473 / 0.046 0.647 / 0.326 / 0.169 0.628
leo-1.7b-v4 0.557 / 0.473 / 0.070 0.826 / 0.684 / 0.141 0.479 / 0.496 / 0.031 0.644 / 0.324 / 0.136 0.627
Jev 1.13 (live API) 0.587 / 0.503 / 0.280 0.790 / 0.693 / 0.064 0.669 / 0.627 / 0.168 0.710 / 0.385 / 0.156 0.689

Multilingual (blind, evaluation only)

Belebele reading comprehension, MMMLU (professionally translated MMLU) and INCLUDE (native regional exams): 50 items per language and suite, one choice request each. None were used for training; Belebele passages that share text with the SIB-200 training data were dropped.

system Belebele MMMLU (knowledge) INCLUDE (knowledge) all ECE
leo-1.7b-v3 0.692 0.479 0.491 0.560 0.101
leo-1.7b-v4 0.680 0.471 0.503 0.557 0.133
Jev 1.13 (live API) 0.917 0.861 0.775 0.857 0.013
belebele accuracy by language
language leo-1.7b-v3 leo-1.7b-v4 Jev 1.13 (live API)
Arabic 0.64 0.62 0.92
Bengali 0.64 0.60 0.92
Chinese 0.84 0.82 0.96
English 0.90 0.90 0.94
French 0.82 0.78 0.94
German 0.82 0.84 0.94
Hindi 0.62 0.62 0.84
Indonesian 0.70 0.72 0.94
Italian 0.74 0.70 0.92
Japanese 0.74 0.72 0.94
Korean 0.70 0.70 0.96
Portuguese 0.80 0.76 0.94
Spanish 0.78 0.70 0.94
Swahili 0.44 0.42 0.92
Yoruba 0.20 0.30 0.74
mmmlu accuracy by language
language leo-1.7b-v3 leo-1.7b-v4 Jev 1.13 (live API)
Arabic 0.48 0.44 0.82
Bengali 0.36 0.32 0.84
Chinese 0.58 0.52 0.90
French 0.50 0.56 0.90
German 0.54 0.50 0.86
Hindi 0.42 0.40 0.80
Indonesian 0.58 0.52 0.92
Italian 0.50 0.52 0.94
Japanese 0.44 0.42 0.90
Korean 0.38 0.38 0.86
Portuguese 0.56 0.56 0.94
Spanish 0.62 0.56 0.96
Swahili 0.36 0.34 0.76
Yoruba 0.38 0.56 0.66
include accuracy by language
language leo-1.7b-v3 leo-1.7b-v4 Jev 1.13 (live API)
Arabic 0.41 0.33 0.71
Bengali 0.42 0.44 0.66
Chinese 0.70 0.76 0.92
French 0.44 0.46 0.78
German 0.48 0.52 0.58
Hindi 0.40 0.48 0.76
Indonesian 0.60 0.62 0.94
Italian 0.54 0.50 0.90
Japanese 0.52 0.58 0.92
Korean 0.38 0.40 0.70
Portuguese 0.46 0.46 0.72
Spanish 0.54 0.48 0.70

Robustness and speed (leo-1.7b-v3)

Option-order sensitivity (how often the top answer changes when the options are shuffled), whether questions sharing a request influence each other, and latency. Leo was timed on one NVIDIA GeForce RTX 3050 (8 GB laptop GPU). Jev's model time is the server time TypeSafe reports; its end-to-end time includes the network round trip from this machine. "2 views" is --order-views 2.

probe leo-1.7b-v3 bf16 leo-1.7b-v3 bf16 2 views Jev (live)
option-order flip rate, emotion (300 rows x 5 shuffles) 0.100 0.055 0.018
mean total-variation shift, emotion 0.063 0.032 0.036
option-order flip rate, fin_topic (300 rows x 5 shuffles) 0.217 0.116 0.069
mean total-variation shift, fin_topic 0.141 0.058 0.063
4 questions together vs one at a time: max prob. difference (100 states) 0.0236 0.0195 0.1300
... top answers that changed 2 3 1
model p50 / p95 ms, short state, 1 question(s) 37 / 46 36 / 41 56 / 84
end-to-end p50 / p95 ms incl. network, short state, 1 question(s) - - 324 / 365
model p50 / p95 ms, short state, 10 question(s) 91 / 92 142 / 144 64 / 83
end-to-end p50 / p95 ms incl. network, short state, 10 question(s) - - 336 / 357
model p50 / p95 ms, short state, 50 question(s) 392 / 393 1082 / 1086 78 / 133
end-to-end p50 / p95 ms incl. network, short state, 50 question(s) - - 358 / 439
model p50 / p95 ms, long state, 1 question(s) 86 / 87 91 / 92 60 / 105
end-to-end p50 / p95 ms incl. network, long state, 1 question(s) - - 333 / 373
model p50 / p95 ms, long state, 10 question(s) 142 / 142 186 / 187 62 / 113
end-to-end p50 / p95 ms incl. network, long state, 10 question(s) - - 336 / 415
model p50 / p95 ms, long state, 50 question(s) 460 / 461 1082 / 1085 72 / 110
end-to-end p50 / p95 ms incl. network, long state, 50 question(s) - - 351 / 382

Limitations and known flaws

Please read these before relying on Leo.

  • False DONE in browser agents. On Google Flights, leo-1.7b-v3 stopped with DONE in 3 of 3 runs without completing the search. It can declare a task done when its action history looks complete even if the page shows no evidence (an open dropdown, a blank render). Always verify outcomes independently, as jev-ultrafast recommends, and never let a DONE trigger anything irreversible. The v4 experiment above targets this.
  • Hard reasoning. JevBench hard tier: 0.396 vs Jev 0.721. Weakest on ambiguous items, long policies with exceptions, trade-offs and multi-hop lookups.
  • World knowledge. MMMLU 0.479 and INCLUDE 0.491, far below Jev. Do not use it as a knowledge source; put the facts it needs into the state.
  • Low-resource languages. Accuracy drops steeply outside the major languages (Yoruba and Swahili are near chance on several suites). Test on your language before deploying.
  • Option order. The top answer changes more often than Jev's when options are shuffled, especially with many similar options. Use --order-views 2 when that matters.
  • Catch-all labels such as "other" or "no emotion" are picked less often than they are right (low macro-F1 on daily_dialog).
  • Calibration away from home. Dev ECE is about 0.01; held-out ECE ranges from about 0.03 to 0.17. Re-fit the temperatures on a few hundred labelled examples of your own traffic before thresholding on probabilities.
  • score questions are the weakest type (dev accuracy about 0.65); treat the expected score as a soft signal.
  • Many questions per request cost more time on Leo than on Jev's servers (see the latency table).
  • Long states beyond the 2,048 training tokens work but are less tested.

Intended use

Good fits: ticket and message routing, intent and topic classification, moderation triage, policy checks with the policy in the state, yes/no checks over documents, rubric grading, next-action choices for agents that verify outcomes, research on calibrated decision models, and a local, private backend for /v1/systemone clients.

Out of scope: decisions about people's health, legal status, credit, housing or employment without qualified human review; sole-filter safety moderation; questions that need world knowledge the state does not contain; unattended agents that can spend money or change accounts.

Training

  • leo-1.7b-v3: 84,315 requests (130,286 questions), 2.0 epochs, 6,556 optimizer steps, learning rate 0.0002, from Qwen/Qwen3-1.7B-Base.
  • leo-1.7b-v4: warm start from leo-1.7b-v3; 39,000 requests (30,000 replayed from v3's training set, 9,000 new evidence-family steps), 1 epoch, learning rate 1e-4; released checkpoint = step 1,200 of 1,911, chosen by calibrated dev loss.
  • fp16 mixed precision, data-parallel on 2x NVIDIA T4 (Kaggle). States cut at 2048 tokens, packed rows at 3072 tokens. About 31 T4 GPU-hours.
  • Option order shuffled; option keys and descriptions varied (bare, described, snake_case, letters, numbers); none/other options appear both when right and when wrong, so they are not a shortcut.

Training data

Human-labelled public datasets (check each licence before commercial use; some are unknown or custom):

source dataset licence requests
ag_news fancyzhx/ag_news unknown 2,500
dbpedia_14 fancyzhx/dbpedia_14 CC BY-SA 3.0 2,000
yahoo_answers community-datasets/yahoo_answers_topics unknown 2,500
banking77 legacy-datasets/banking77 CC BY 4.0 2,500
clinc_oos clinc/clinc_oos CC BY 3.0 2,500
snli stanfordnlp/snli CC BY-SA 4.0 2,000
multi_nli nyu-mll/multi_nli mixed (see dataset card) 2,500
mrpc nyu-mll/glue other (GLUE / MRPC terms) 1,000
boolq google/boolq CC BY-SA 3.0 2,000
arc_easy allenai/ai2_arc CC BY-SA 4.0 1,200
arc_challenge allenai/ai2_arc CC BY-SA 4.0 800
commonsense_qa tau/commonsense_qa MIT 2,000
imdb stanfordnlp/imdb other 1,500
yelp Yelp/yelp_review_full other (Yelp dataset terms) 2,500
sms_spam ucirvine/sms_spam unknown on card 1,200
prompt_injections deepset/prompt-injections Apache-2.0 546
jailbreak jackhhao/jailbreak-classification Apache-2.0 900
helpsteer2 nvidia/HelpSteer2 CC BY 4.0 2,000
massive AmazonScience/massive CC BY 4.0 11,146
sib200 Davlan/sib200 CC BY-SA 4.0 6,384
ml_sentiment tyqiangz/multilingual-sentiments Apache-2.0 4,607
mind2web osunlp/Mind2Web CC BY 4.0 7,132

Code-generated families (labels computed by code, not by a model):

family what it teaches requests
syn_browser browser-agent episodes on simulated sites, in jev-ultrafast's exact request format 12,000
syn_field_reference questions that point at a JSON field by path 1,800
syn_policy policy compliance checks 1,800
syn_unknowable questions the state cannot answer 500
syn_conversation multi-turn conversations 900
syn_multi_hop multi-hop lookups across records 1,500
syn_long_state a relevant fact inside long irrelevant state 1,000
syn_injection instructions injected into state 700
syn_numeric counting, arithmetic and date comparison 1,500
syn_policy_exceptions policies with exceptions 1,200
syn_browser_evidence (v4 only) browser screens where DONE needs visible evidence (overlays, pop-overs, blank and loading renders, stale or unsubmitted results, values a site does not offer) 9,000

Never used for training: any output of TypeSafe's models, the four held-out datasets, Belebele, MMMLU, INCLUDE, JevBench, and the sites and goals of the browser benchmark.

Versions

version branch notes
leo-1.7b-v3 main recommended
leo-1.7b-v4 v4-experimental honest BLOCKED instead of false DONE; fails open-an-item tasks; lower JevBench
leo-0.6b-v0 / v1 not released Qwen3-0.6B prototypes

Licence

Weights and code: Apache-2.0; the Qwen3 base is Apache-2.0 too. The training mix includes datasets under CC BY-SA, custom and unknown terms (listed above). For a clean licence chain, retrain on the permissive subset with the pipeline on GitHub.

Citation

@misc{leo2026,
  title  = {Leo: an open-weight decision model with calibrated typed answers},
  author = {Baranwal, Suparva},
  year   = {2026},
  url    = {https://huggingface.co/Suparva/leo-1.7b},
  note   = {Code: https://github.com/SuparvaCode/leo}
}

Acknowledgements

Qwen for the base model; browser-use/jev-ultrafast, fstandhartinger/jevbench and elcronos/jev-vs-open-decision-models for the benchmarks; TypeSafe for documenting the /v1/systemone format publicly.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Suparva/leo-1.7b

Adapter
(76)
this model

Datasets used to train Suparva/leo-1.7b