Free trial API: https://leo.kognare.com (key: free). Code: https://github.com/SuparvaCode/leo. On Hugging Face the base model is not bundled; Leo.load downloads Qwen/Qwen3-4B-Base at the pinned revision.

Leo-1 (4B)

Leo's final model from the v5 round: an open-weight decision model. Send a state (text or JSON) and typed questions (choice, score, noul yes/no); get calibrated probabilities from one forward pass. It speaks the POST /v1/systemone wire format, so TypeSafe/Jev clients (for example naturalcodz) work by changing the base URL.

Internally this is leo-4b-soup3: an exact weight merge of leo-4b-v5 (70%) and leo-4b-v5.1 (30%) on Qwen/Qwen3-4B-Base (revision 906bfd4b, bundled in base/), with a LoRA of rank 64, re-calibrated on the dev set (dev ECE 0.007).

What is in this folder

path what
adapter/, leo_head.safetensors, leo_config.json Leo-1's weights and calibration
base/ the exact Qwen3-4B-Base weights it was trained on (Apache-2.0), so nothing is downloaded
leo/ inference code and the HTTP server
results/ the benchmark files behind the numbers below
verify.py offline self-check: re-answers the blind short-input suite and compares with the recorded results
SHA256SUMS.txt checksums of every file

Total size is about 8 GB. Running it needs about 9 GB of GPU memory in bf16 (for example an L4, A10, RTX 3090/4090 or better), or a CPU with about 12 GB free RAM (slow: about 3–4 s per request).

Use

pip install -r requirements.txt
python verify.py                      # prints OK when the files are intact
import sys; sys.path.insert(0, r"D:\Leo-1")
from leo.infer import Leo
leo = Leo.load(r"D:\Leo-1", dtype="bf16")        # device="cpu" if the GPU is too small
print(leo.system_one("I was charged twice this month.",
      {"team": {"type": "choice", "instructions": "Which team?", "criteria": {"billing": None, "technical": None, "other": None}},
       "angry": {"type": "noul", "instructions": "angry"}}))

HTTP server (127.0.0.1 only unless a key is set):

cd D:\Leo-1
set LEO_API_KEY=change-me
python -m leo.serve --model . --port 8000 --canonicalize input

--canonicalize input rewrites bare yes/no conditions ("angry") into questions, which raised the short-input score from 0.953 to 0.972. States longer than 8,192 tokens get a 422 by default instead of being silently cut.

Results (identical requests; Jev = live jev-1.13.0, September 2026)

benchmark Leo-1 Jev 1.13 released leo-1.7b-v3
Browser tasks, jev-ultrafast (21 runs) 18/21 18/21 18/21
naturalcodz drop-in (37 scenarios) 35/37 36/37 32/37
Blind short-input suite (211 conditions) 0.953 (0.972 with rewrite) 0.976 0.872
Short-input calibration error (lower is better) 0.029 0.075 0.093
JevBench all / standard / hard 0.701 / 0.944 / 0.414 0.861 / 0.986 / 0.721 0.697 / 0.958 / 0.396
Held-out classification, mean accuracy 0.650 0.689 0.628
tweet_topic 0.840 0.790 0.845
Held-out calibration error (lower is better) 0.069 0.167 0.099
Multilingual (Belebele / MMMLU / INCLUDE) 0.681 / 0.481 / 0.576 0.917 / 0.861 / 0.775 0.692 / 0.479 / 0.491
False DONE on held-out browser screens 0.000 not measured 0.067
Option-order flips, emotion / fin_topic 0.074 / 0.101 0.018 / 0.069 0.100 / 0.217

Where it stands

  • Level with Jev or close: browser agents (18/21), naturalcodz-style calls, short and one-word conditions, spam / PII / toxicity checks, routing, JevBench easy and standard. Its probabilities are better calibrated than Jev's on the short-input and held-out suites.
  • Clearly behind Jev: hard reasoning (JevBench hard 0.41 vs 0.72: multi-hop, dates and numbers, ambiguous items), world knowledge (MMMLU 0.48 vs 0.86) and low-resource languages. It is not a Jev replacement for those.
  • Known flaws: Google Flights still fails (it no longer claims DONE, but it does not finish the search); a refund message was routed to "sales" in naturalcodz; isSafe misses one insult (Jev misses the same one); option order still moves answers more than Jev (use --order-views 2).
  • Always verify agent outcomes independently; never let a DONE trigger something irreversible.

Not trained on any Jev output: Jev was called only to score it. Weights and code are Apache-2.0; training datasets keep their own licences (see the Leo repository).

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Suparva/leo-1

Adapter
(95)
this model