__ __ / /___ _ ___ __ __ ___ / // // _ `// _ \/ // /(_-< \___/ \_,_//_//_/\_,_//___/
Janus 0.8B: calibrated typed decisions in one forward pass.
Janus is an independent reconstruction of a "System 1" decision model, built from public descriptions of the
interface. Janus 0.8B is the small model: give it a state (a ticket, a log, a policy, a JSON record) and any number of choice,
score and noul (yes/no) questions. In one forward pass, it returns a calibrated probability for every option
of every question, with no text generation. It is a LoRA adapter (r64) plus a pointer decision head on
Qwen/Qwen3.5-0.8B at revision 2fc06364715b967f1860aea9cf38778875588b17.
It needs about 2 GB of GPU memory and also runs on CPU. For harder decisions, use
Janus 4B. The state is encoded once however many questions ride on it, each
question is answered independently, a choice can have up to 255 options, and one short question takes about 5 ms on
an RTX 5090. It is trained on data in 51 locales.
Code, server and Docker image: github.com/IcarusAICo/janus.
| Questions | choice (1 to 255 options), score (2 to 10 ordered levels, with the expected level), noul (p(true)) |
| Input | Up to 16,384 tokens in the Docker image (--max-tokens sets it); trained on requests up to 9,216 |
| Wire format | The POST /v1/systemone request and response shape; output tokens are always 0 |
| Memory | bf16 on an RTX 3090: 1.8 GB peak with a 150-token state, 2.2 GB with 4k tokens; about 4 GB more with the server's start-up pre-capture |
Results
All numbers are from 2026-09-23/24, with every system served by its own code: accuracy on one RTX 3090, latency on
an RTX 3090 and an RTX 5090. Janus 4B is the
v3-instruct checkpoint. Full tables and sources: release/RESULTS.md in the GitHub repository.
| set | metric | Janus 0.8B | Janus 4B | JevK5 v0.2 | Laya 0.3.11 |
|---|---|---|---|---|---|
| JevBench public items, easy / standard / hard | accuracy | 1.000 / 0.778 / 0.613 | 1.000 / 0.972 / 0.712 | 1.000 / 0.972 / 0.730 | 0.958 / 0.694 / 0.333 |
| JevBench public items, all 231 | accuracy | 0.745 | 0.853 | 0.861 | 0.576 |
| JevBench public hard tier | ECE | 0.072 | 0.071 | 0.062 | 0.206 |
| MASSIVE, 51 locales, official test split | 20-way intent accuracy / ECE | 0.635 / 0.059 | 0.767 / 0.039 | refused (16-option limit) | 0.272 / 0.173 |
| XNLI, 15 languages, official test split | accuracy / ECE | 0.655 / 0.044 | 0.757 / 0.055 | 0.608 / 0.165 | 0.522 / 0.051 |
The JevBench rows are our runs of the 231 public items through JevBench's own runner. They are not an official JevBench score (the official v1.4 score also pools 308 sealed items), and they were never used to select a checkpoint. A v1.4 what-if with an assumed sealed accuracy of 0.31 for every system (not an official score), scored the way the v1.4.2 board scores self-hosted rows (list-price cost, timing on an RTX 5090), puts Janus 0.8B at 43.3. Under our earlier method (GPU-hour cost, RTX 3090 timing), which also scored Laya, it was 45.8 against Laya's 23.4.
Latency. Median / p90 ms per request, warm, one request at a time, on the final serving code (start-up pre-capture on). Every system was served by its own code on the same card. Laya truncates long states to its context (26-35% of long-state requests), which also shortens its latency there. On the smaller cards Janus's long-state p90 includes graph re-captures after a memory release.
| request | card | Janus 0.8B | Janus 4B | JevK5 v0.2 | Laya 0.3.11 |
|---|---|---|---|---|---|
| 1 short question (XNLI) | RTX 5090 | 5.0 / 5.4 | 17.3 / 18.1 | 18.7 / 19.6 | 6.1 / 7.4 |
| RTX 3090 | 10.7 / 11.2 | 47.0 / 47.5 | 40.9 / 41.4 | 6.3 / 8.9 | |
| public-dataset test (1-2 questions) | RTX 5090 | 8.7 / 14.4 | 27.9 / 55.7 | 40.3 / 117.8 | 9.1 / 14.8 |
| RTX 3090 | 16.9 / 32.6 | 69.7 / 149.4 | 82.0 / 322.1 | 14.5 / 31.4 | |
| 1 question, long state (hard-tier v2) | RTX 5090 | 7.9 / 44.4 | 27.0 / 161.6 | 32.6 / 247.0 | 9.1 / 13.9 |
| RTX 3090 | 16.4 / 94.3 | 75.0 / 462.1 | 80.6 / 647.7 | 14.9 / 23.8 | |
| 8 questions, one long state | RTX 5090 | 34.3 / 88.5 | 107.0 / 271.4 | 251.8 / 1,964.4 | 26.0 / 56.4 |
| RTX 3090 | 73.4 / 169.6 | 337.8 / 857.3 | 641.7 / 5,194.8 | 66.6 / 110.9 |
Use
pip install "janus @ git+https://github.com/IcarusAICo/janus"
import janus
m = janus.load("cmxu/janus-0.8b", device="cuda") # downloads this repo and the Qwen/Qwen3.5-0.8B backbone
response = m.predict(
"Refunds need a receipt and a purchase within 30 days. The customer bought 12 days ago and has no receipt.",
{"refund": {"type": "noul", "instructions": "Is a refund permitted under the policy?"},
"route": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "Payments and refunds", "support": "Anything else"}}},
)
response["answers"]["refund"]["noul"] # p(true)
response["answers"]["route"]["probabilities"] # {"billing": ..., "support": ...}
m.predict_batch([request, ...]) scores many /v1/systemone bodies with shared forward passes. To serve over HTTP
from a download of this repo:
python -m janus serve --checkpoint model.pt --calibration calibration.json --model-id janus-0.8b --device cuda
The repository's Dockerfile builds an offline container around a pinned revision of this repo.
Files
| file | contents |
|---|---|
model.pt |
trainable tensors only (LoRA + decision head), sha256 321eb9ec4c888fe0e2ea73d56bbf240f9fc32618dc2112d9690e37811a2a5e22; the backbone is fetched from Qwen/Qwen3.5-0.8B at 2fc06364715b967f1860aea9cf38778875588b17 |
calibration.json |
global, per-family and per-option-count temperatures, bound to model.pt by its sha256 |
janus_config.json |
model id and pinned base model |
The global temperature is 0.7290. A request whose group_id starts with family: gets that family's temperature.
Calibrated families: adversarial, ambiguous, arena, badges, civil, delivery, evidence, helpsteer,
judge_hard, long_policy, massive, multi_hop, ord, post, probability, quality, record_match, rel,
retrieval, routing_hard, tabfact, temporal_numeric, tradeoff, trap, urgency, wikispeedia.
Training
Run: runs/publish/qwen35_08b_distill_v2, config configs/publish/qwen35_08b_distill_v2.json.
Data. 57,304 requests (
data/distill-v2):- synthetic decision families, some generated by code and some written by GPT-5.6 Luna and reviewed by GPT-5.6 Terra (hard-tier policies, multi-step lookups, dates and numbers, probability, trade-offs, answer judging);
- public datasets converted to typed requests (the
datasetsabove); - Wikispeedia next-click decisions, scripted browser-action rows, long-context and many-option rows;
- 4,080 multilingual rows from MASSIVE's train split (51 locales). No XNLI rows, so its XNLI-15 score is zero-shot.
Rows are deduplicated by state and kept disjoint from evaluation states. The hard-tier generators reject any state sharing a 12-word sequence with a public JevBench item. The datasets and their licences are listed in the GitHub README.
Distillation. Janus 35B-A3B labelled every non-multilingual request. A one-hot gold target became 0.5 gold + 0.5 teacher distribution.
Schedule. One epoch from scratch: 1,791 steps of 32 requests, 50 warm-up steps, then cosine decay (backbone lr 2e-4, head lr 1e-3). Cross-entropy with label smoothing 0.1 plus an option-order consistency term (weight 0.1). Local RTX 3090, about 5.2 h.
Selection. Step 1,500 of 1,791, by the rule fixed before training: dev NLL plus a penalty for confident errors on held-out families. Dev NLL 0.464, accuracy 0.822. The held-out NLL on unseen families, 1.007 against 0.896 for a uniform guess, carries a small penalty. JevBench items were never used for selection.
Calibration. Temperatures fitted after selection on 3,307 held-out requests that share no state or group with training or selection data. We also tested a hard-tier refit of the global temperature. We rejected it because it hurt judge v2 (ECE 0.049 to 0.115) and public-dataset test calibration (0.028 to 0.041).
License and credits
Code and adapter weights: Apache-2.0. The base model Qwen3.5-0.8B and the teacher's backbone Qwen3.6-35B-A3B are by the Qwen team (Apache-2.0). The training datasets keep their own licences.
"Jev", "System One" and TypeSafe are names of TypeSafe AI's products, used here only to describe interface compatibility.