Instructions to use Suparva/leo-1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Suparva/leo-1 with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Free trial API: https://leo.kognare.com (key:
free). Code: https://github.com/SuparvaCode/leo. On Hugging Face the base model is not bundled;Leo.loaddownloadsQwen/Qwen3-4B-Baseat the pinned revision.
Leo-1 (4B)
Leo's final model from the v5 round: an open-weight decision model. Send a state (text or JSON) and typed questions (choice, score, noul yes/no); get calibrated probabilities from one forward pass. It speaks the POST /v1/systemone wire format, so TypeSafe/Jev clients (for example naturalcodz) work by changing the base URL.
Internally this is leo-4b-soup3: an exact weight merge of leo-4b-v5 (70%) and leo-4b-v5.1 (30%) on Qwen/Qwen3-4B-Base (revision 906bfd4b, bundled in base/), with a LoRA of rank 64, re-calibrated on the dev set (dev ECE 0.007).
What is in this folder
| path | what |
|---|---|
adapter/, leo_head.safetensors, leo_config.json |
Leo-1's weights and calibration |
base/ |
the exact Qwen3-4B-Base weights it was trained on (Apache-2.0), so nothing is downloaded |
leo/ |
inference code and the HTTP server |
results/ |
the benchmark files behind the numbers below |
verify.py |
offline self-check: re-answers the blind short-input suite and compares with the recorded results |
SHA256SUMS.txt |
checksums of every file |
Total size is about 8 GB. Running it needs about 9 GB of GPU memory in bf16 (for example an L4, A10, RTX 3090/4090 or better), or a CPU with about 12 GB free RAM (slow: about 3–4 s per request).
Use
pip install -r requirements.txt
python verify.py # prints OK when the files are intact
import sys; sys.path.insert(0, r"D:\Leo-1")
from leo.infer import Leo
leo = Leo.load(r"D:\Leo-1", dtype="bf16") # device="cpu" if the GPU is too small
print(leo.system_one("I was charged twice this month.",
{"team": {"type": "choice", "instructions": "Which team?", "criteria": {"billing": None, "technical": None, "other": None}},
"angry": {"type": "noul", "instructions": "angry"}}))
HTTP server (127.0.0.1 only unless a key is set):
cd D:\Leo-1
set LEO_API_KEY=change-me
python -m leo.serve --model . --port 8000 --canonicalize input
--canonicalize input rewrites bare yes/no conditions ("angry") into questions, which raised the short-input score from 0.953 to 0.972. States longer than 8,192 tokens get a 422 by default instead of being silently cut.
Results (identical requests; Jev = live jev-1.13.0, September 2026)
| benchmark | Leo-1 | Jev 1.13 | released leo-1.7b-v3 |
|---|---|---|---|
| Browser tasks, jev-ultrafast (21 runs) | 18/21 | 18/21 | 18/21 |
| naturalcodz drop-in (37 scenarios) | 35/37 | 36/37 | 32/37 |
| Blind short-input suite (211 conditions) | 0.953 (0.972 with rewrite) | 0.976 | 0.872 |
| Short-input calibration error (lower is better) | 0.029 | 0.075 | 0.093 |
| JevBench all / standard / hard | 0.701 / 0.944 / 0.414 | 0.861 / 0.986 / 0.721 | 0.697 / 0.958 / 0.396 |
| Held-out classification, mean accuracy | 0.650 | 0.689 | 0.628 |
| tweet_topic | 0.840 | 0.790 | 0.845 |
| Held-out calibration error (lower is better) | 0.069 | 0.167 | 0.099 |
| Multilingual (Belebele / MMMLU / INCLUDE) | 0.681 / 0.481 / 0.576 | 0.917 / 0.861 / 0.775 | 0.692 / 0.479 / 0.491 |
| False DONE on held-out browser screens | 0.000 | not measured | 0.067 |
| Option-order flips, emotion / fin_topic | 0.074 / 0.101 | 0.018 / 0.069 | 0.100 / 0.217 |
Where it stands
- Level with Jev or close: browser agents (18/21), naturalcodz-style calls, short and one-word conditions, spam / PII / toxicity checks, routing, JevBench easy and standard. Its probabilities are better calibrated than Jev's on the short-input and held-out suites.
- Clearly behind Jev: hard reasoning (JevBench hard 0.41 vs 0.72: multi-hop, dates and numbers, ambiguous items), world knowledge (MMMLU 0.48 vs 0.86) and low-resource languages. It is not a Jev replacement for those.
- Known flaws: Google Flights still fails (it no longer claims DONE, but it does not finish the search); a refund message was routed to "sales" in naturalcodz;
isSafemisses one insult (Jev misses the same one); option order still moves answers more than Jev (use--order-views 2). - Always verify agent outcomes independently; never let a DONE trigger something irreversible.
Not trained on any Jev output: Jev was called only to score it. Weights and code are Apache-2.0; training datasets keep their own licences (see the Leo repository).
- Downloads last month
- -
Model tree for Suparva/leo-1
Base model
Qwen/Qwen3-4B-Base