__                    
 __ / /___ _ ___  __ __ ___
/ // // _ `// _ \/ // /(_-<
\___/ \_,_//_//_/\_,_//___/

Janus 0.8B: calibrated typed decisions in one forward pass.

Janus is an independent reconstruction of a "System 1" decision model, built from public descriptions of the interface. Janus 0.8B is the small model: give it a state (a ticket, a log, a policy, a JSON record) and any number of choice, score and noul (yes/no) questions. In one forward pass, it returns a calibrated probability for every option of every question, with no text generation. It is a LoRA adapter (r64) plus a pointer decision head on Qwen/Qwen3.5-0.8B at revision 2fc06364715b967f1860aea9cf38778875588b17. It needs about 2 GB of GPU memory and also runs on CPU. For harder decisions, use Janus 4B. The state is encoded once however many questions ride on it, each question is answered independently, a choice can have up to 255 options, and one short question takes about 5 ms on an RTX 5090. It is trained on data in 51 locales.

Code, server and Docker image: github.com/IcarusAICo/janus.

Questions choice (1 to 255 options), score (2 to 10 ordered levels, with the expected level), noul (p(true))
Input Up to 16,384 tokens in the Docker image (--max-tokens sets it); trained on requests up to 9,216
Wire format The POST /v1/systemone request and response shape; output tokens are always 0
Memory bf16 on an RTX 3090: 1.8 GB peak with a 150-token state, 2.2 GB with 4k tokens; about 4 GB more with the server's start-up pre-capture

Results

All numbers are from 2026-09-23/24, with every system served by its own code: accuracy on one RTX 3090, latency on an RTX 3090 and an RTX 5090. Janus 4B is the v3-instruct checkpoint. Full tables and sources: release/RESULTS.md in the GitHub repository.

set metric Janus 0.8B Janus 4B JevK5 v0.2 Laya 0.3.11
JevBench public items, easy / standard / hard accuracy 1.000 / 0.778 / 0.613 1.000 / 0.972 / 0.712 1.000 / 0.972 / 0.730 0.958 / 0.694 / 0.333
JevBench public items, all 231 accuracy 0.745 0.853 0.861 0.576
JevBench public hard tier ECE 0.072 0.071 0.062 0.206
MASSIVE, 51 locales, official test split 20-way intent accuracy / ECE 0.635 / 0.059 0.767 / 0.039 refused (16-option limit) 0.272 / 0.173
XNLI, 15 languages, official test split accuracy / ECE 0.655 / 0.044 0.757 / 0.055 0.608 / 0.165 0.522 / 0.051

The JevBench rows are our runs of the 231 public items through JevBench's own runner. They are not an official JevBench score (the official v1.4 score also pools 308 sealed items), and they were never used to select a checkpoint. A v1.4 what-if with an assumed sealed accuracy of 0.31 for every system (not an official score), scored the way the v1.4.2 board scores self-hosted rows (list-price cost, timing on an RTX 5090), puts Janus 0.8B at 43.3. Under our earlier method (GPU-hour cost, RTX 3090 timing), which also scored Laya, it was 45.8 against Laya's 23.4.

Latency. Median / p90 ms per request, warm, one request at a time, on the final serving code (start-up pre-capture on). Every system was served by its own code on the same card. Laya truncates long states to its context (26-35% of long-state requests), which also shortens its latency there. On the smaller cards Janus's long-state p90 includes graph re-captures after a memory release.

request card Janus 0.8B Janus 4B JevK5 v0.2 Laya 0.3.11
1 short question (XNLI) RTX 5090 5.0 / 5.4 17.3 / 18.1 18.7 / 19.6 6.1 / 7.4
RTX 3090 10.7 / 11.2 47.0 / 47.5 40.9 / 41.4 6.3 / 8.9
public-dataset test (1-2 questions) RTX 5090 8.7 / 14.4 27.9 / 55.7 40.3 / 117.8 9.1 / 14.8
RTX 3090 16.9 / 32.6 69.7 / 149.4 82.0 / 322.1 14.5 / 31.4
1 question, long state (hard-tier v2) RTX 5090 7.9 / 44.4 27.0 / 161.6 32.6 / 247.0 9.1 / 13.9
RTX 3090 16.4 / 94.3 75.0 / 462.1 80.6 / 647.7 14.9 / 23.8
8 questions, one long state RTX 5090 34.3 / 88.5 107.0 / 271.4 251.8 / 1,964.4 26.0 / 56.4
RTX 3090 73.4 / 169.6 337.8 / 857.3 641.7 / 5,194.8 66.6 / 110.9

Use

pip install "janus @ git+https://github.com/IcarusAICo/janus"
import janus

m = janus.load("cmxu/janus-0.8b", device="cuda")   # downloads this repo and the Qwen/Qwen3.5-0.8B backbone
response = m.predict(
    "Refunds need a receipt and a purchase within 30 days. The customer bought 12 days ago and has no receipt.",
    {"refund": {"type": "noul", "instructions": "Is a refund permitted under the policy?"},
     "route": {"type": "choice", "instructions": "Which team should handle this?",
               "criteria": {"billing": "Payments and refunds", "support": "Anything else"}}},
)
response["answers"]["refund"]["noul"]           # p(true)
response["answers"]["route"]["probabilities"]   # {"billing": ..., "support": ...}

m.predict_batch([request, ...]) scores many /v1/systemone bodies with shared forward passes. To serve over HTTP from a download of this repo:

python -m janus serve --checkpoint model.pt --calibration calibration.json --model-id janus-0.8b --device cuda

The repository's Dockerfile builds an offline container around a pinned revision of this repo.

Files

file contents
model.pt trainable tensors only (LoRA + decision head), sha256 321eb9ec4c888fe0e2ea73d56bbf240f9fc32618dc2112d9690e37811a2a5e22; the backbone is fetched from Qwen/Qwen3.5-0.8B at 2fc06364715b967f1860aea9cf38778875588b17
calibration.json global, per-family and per-option-count temperatures, bound to model.pt by its sha256
janus_config.json model id and pinned base model

The global temperature is 0.7290. A request whose group_id starts with family: gets that family's temperature. Calibrated families: adversarial, ambiguous, arena, badges, civil, delivery, evidence, helpsteer, judge_hard, long_policy, massive, multi_hop, ord, post, probability, quality, record_match, rel, retrieval, routing_hard, tabfact, temporal_numeric, tradeoff, trap, urgency, wikispeedia.

Training

Run: runs/publish/qwen35_08b_distill_v2, config configs/publish/qwen35_08b_distill_v2.json.

  • Data. 57,304 requests (data/distill-v2):

    • synthetic decision families, some generated by code and some written by GPT-5.6 Luna and reviewed by GPT-5.6 Terra (hard-tier policies, multi-step lookups, dates and numbers, probability, trade-offs, answer judging);
    • public datasets converted to typed requests (the datasets above);
    • Wikispeedia next-click decisions, scripted browser-action rows, long-context and many-option rows;
    • 4,080 multilingual rows from MASSIVE's train split (51 locales). No XNLI rows, so its XNLI-15 score is zero-shot.

    Rows are deduplicated by state and kept disjoint from evaluation states. The hard-tier generators reject any state sharing a 12-word sequence with a public JevBench item. The datasets and their licences are listed in the GitHub README.

  • Distillation. Janus 35B-A3B labelled every non-multilingual request. A one-hot gold target became 0.5 gold + 0.5 teacher distribution.

  • Schedule. One epoch from scratch: 1,791 steps of 32 requests, 50 warm-up steps, then cosine decay (backbone lr 2e-4, head lr 1e-3). Cross-entropy with label smoothing 0.1 plus an option-order consistency term (weight 0.1). Local RTX 3090, about 5.2 h.

  • Selection. Step 1,500 of 1,791, by the rule fixed before training: dev NLL plus a penalty for confident errors on held-out families. Dev NLL 0.464, accuracy 0.822. The held-out NLL on unseen families, 1.007 against 0.896 for a uniform guess, carries a small penalty. JevBench items were never used for selection.

  • Calibration. Temperatures fitted after selection on 3,307 held-out requests that share no state or group with training or selection data. We also tested a hard-tier refit of the global temperature. We rejected it because it hurt judge v2 (ECE 0.049 to 0.115) and public-dataset test calibration (0.028 to 0.041).

License and credits

Code and adapter weights: Apache-2.0. The base model Qwen3.5-0.8B and the teacher's backbone Qwen3.6-35B-A3B are by the Qwen team (Apache-2.0). The training datasets keep their own licences.

"Jev", "System One" and TypeSafe are names of TypeSafe AI's products, used here only to describe interface compatibility.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for cmxu/janus-0.8b

Adapter
(267)
this model

Datasets used to train cmxu/janus-0.8b