Weiche-395M: a small local decision model for LLM auto-routing
Weiche (German for a railway switch) reads one user turn and decides which kind of model should handle it. It is ModernBERT-large (395M) with typed decision heads, trained for the open-source auto-model-router. One forward pass returns:
| Output | Meaning |
|---|---|
category |
9-way choice: coding, agentic, math, knowledge, long_context, tool_use, design, summarisation, general |
scalars[1] difficulty |
rubric difficulty, 0 = trivial … 1 = frontier (the router's 0–4 scale ÷ 4) |
scalars[0] outcome difficulty |
share of a model panel expected to fail (learned from measured outcomes) |
scalars[2] stakes |
cost of an unnoticed wrong answer, 0–1 |
noul |
P(needs_tools), P(needs_vision), P(needs_long_context), P(follow_up) |
succ |
P(success) of a small, a mid and a strong model tier, learned from measured outcomes |
route |
experimental: small / mid / strong given 14 numeric router-state features (cache warm/cold, latency budget, value of a correct answer, tier prices and latencies) |
It runs on CPU through ONNX Runtime (no PyTorch needed) and in the browser through onnxruntime-web (WebGPU).
Results
All numbers come from held-out data that was never used for training. The eval scripts
and raw result JSON files are in eval/ and in the
GitHub repo. The baselines are the router's
existing decision paths: its heuristic, the local Laya 421M classifier (backend: local)
and the hosted Jev API (backend: hosted). Jev was called only for this evaluation. Its
outputs were never used for training, calibration or data selection.
The "+ fitted calibration" rows give the trait-only classifiers a per-tier logistic
calibration from their traits to P(success). We fitted it on a separate calibration split,
which is more help than the router's fixed success curve gives them. Weiche's success heads
need no calibration. Routing picks, for each request, argmax P(success_t) − λ·cost_t and
sweeps λ, as the router's expected-cost policy does. APGR is the average share of the
quality gap between always-small and always-strong that is recovered across 20 cost budgets.
| Classifier | Category acc. (held-out, n=1134) | Macro-F1 | Category acc. on the router's own 70 held-out tasks | needs_tools acc. | Difficulty rank corr. |
|---|---|---|---|---|---|
| Weiche-395M (this model) | 88.5 % | 0.872 | 94.9 % (59 tasks) | 98.4 % | 0.739 |
| Weiche-0.6B arm (Qwen3-0.6B, not released) | 88.5 % | 0.873 | 98.3 % (59 tasks) | 99.2 % | 0.736 |
| hosted Jev 1.13 (API) | 85.8 % | 0.843 | 79.7 % (59 tasks) | 96.0 % | 0.753 |
Laya 421M (router local backend) |
9.3 % | 0.112 | 33.9 % (59 tasks) | 80.5 % | 0.297 |
RouterBench held-out (n=1500; tiers Mistral-7B / Mixtral-8x7B / GPT-4-1106)
| Decision path | Avg. gap recovered (APGR) | Quality at 10 % of strong-tier cost | at 25 % | Cost share to recover 80 % of the gap | 95 % of the gap |
|---|---|---|---|---|---|
| Router heuristic (no classifier) | 0.494 | 47.2 % | 47.2 % | 100.0 % | 100.0 % |
Laya 421M (router local backend) + fitted calibration |
0.812 | 48.9 % | 53.8 % | 38.9 % | 79.5 % |
| hosted Jev 1.13 (API) + fitted calibration | 0.850 | 50.3 % | 59.6 % | 28.2 % | 86.2 % |
| Weiche-395M traits + fitted calibration | 0.834 | 48.7 % | 58.2 % | 41.0 % | 88.1 % |
| Weiche-395M success heads | 0.862 | 51.4 % | 59.0 % | 31.4 % | 65.2 % |
| Weiche-0.6B arm success heads | 0.864 | 51.8 % | 59.0 % | 30.4 % | 68.7 % |
Always-small quality 28.4 %, always-mid 47.2 %, always-strong 68.6 %. An oracle that knows every outcome matches always-strong quality at 17.1 % of its cost and tops out at 77.2 %.
Measured router-style tasks finished after the data freeze (n=1103; tiers Mistral-Nemo-12B / Gemma-4-31B / DeepSeek-V3.2)
| Decision path | Avg. gap recovered (APGR) | Quality at 10 % of strong-tier cost | at 25 % | Cost share to recover 80 % of the gap | 95 % of the gap |
|---|---|---|---|---|---|
| Router heuristic (no classifier) | 0.459 | 42.5 % | 42.5 % | 60.8 % | 60.8 % |
Laya 421M (router local backend) + fitted calibration |
0.747 | 48.2 % | 64.1 % | 47.5 % | 56.2 % |
| hosted Jev 1.13 (API) + fitted calibration | 0.734 | 48.9 % | 56.4 % | 50.5 % | 56.9 % |
| Weiche-395M traits + fitted calibration | 0.772 | 52.4 % | 62.4 % | 44.0 % | 54.4 % |
| Weiche-395M success heads | 0.803 | 52.9 % | 72.4 % | 44.5 % | 52.7 % |
| Weiche-0.6B arm success heads | 0.805 | 54.0 % | 71.4 % | 41.9 % | 52.7 % |
Always-small quality 42.5 %, always-mid 95.0 %, always-strong 94.0 %. An oracle that knows every outcome matches always-strong quality at 38.3 % of its cost and tops out at 98.6 %.
| CPU latency per decision, 4 threads, same busy host | p50 | p95 |
|---|---|---|
| Weiche-395M ONNX int4 | 288 ms | 393 ms |
| Weiche-395M ONNX int8 (lossy) | 243 ms | 382 ms |
| Weiche-395M ONNX fp16 | 386 ms | 545 ms |
| Laya 421M (PyTorch, 7 questions) | 11061 ms | 12308 ms |
The difficulty rank correlation compares against the averaged rubric labels of two model families. On that measure Weiche and hosted Jev are close (0.739 vs 0.753).
Use it in auto-model-router
policy:
classifier: {backend: local-route-head, model: benchmarkheaven/weiche-395m, variant: fp16, threads: 2}
variant is fp16 (default, matches fp32 within 1e-3), int4 (423 MB, 96 % category
agreement with fp32) or int8 (lossy for this architecture, 92 % agreement; not
recommended). The first request downloads the chosen ONNX file once.
Use it directly (ONNX Runtime, Python)
import numpy as np, onnxruntime as ort
from huggingface_hub import snapshot_download
from tokenizers import Tokenizer
d = snapshot_download("benchmarkheaven/weiche-395m", allow_patterns=["model_fp16.onnx*", "tokenizer.json", "routerhead.py"])
tok = Tokenizer.from_file(f"{d}/tokenizer.json"); tok.enable_truncation(512)
sess = ort.InferenceSession(f"{d}/model_fp16.onnx", providers=["CPUExecutionProvider"])
def render(request, context=""): # the training-time input format (see routerhead.py)
if len(request) > 2400:
request = request[:1400] + "\n[...]\n" + request[-1000:]
return f"Context: {context.strip()[:600] or '(new conversation)'}\nRequest: {request}"
ids = np.array([tok.encode(render("Fix the flaky test in tests/test_api.py and run the suite",
"3 messages, 2 from the user. Tools: Bash, Read, Edit.")).ids])
category, scalars, noul, succ, route = sess.run(None, {"input_ids": ids, "attention_mask": np.ones_like(ids),
"state": np.zeros((1, 14), np.float32)})
PyTorch: encoder/ holds the fine-tuned ModernBERT (HF format) and heads.safetensors
holds the heads. routerhead.py defines RouterHead and render:
from transformers import AutoModel, AutoTokenizer
from safetensors.torch import load_file
from routerhead import RouterHead, probs, render
enc = AutoModel.from_pretrained("benchmarkheaven/weiche-395m", subfolder="encoder")
model = RouterHead(enc, 1024).eval()
model.load_state_dict(load_file("heads.safetensors"), strict=False)
Browser: model_int4.onnx ran under onnxruntime-web 1.30 with the WebGPU execution provider
(eval/webgpu_smoke_int4.txt). Its outputs matched CPU ONNX Runtime within 1e-4. That check
used Chrome's software WebGPU adapter, so it proves operator coverage, not speed.
model_fp16.onnx needs a GPU with shader-f16; the software adapter has none, so the fp16
browser path is not yet verified on real hardware.
Training data
We did not use any TypeSafe/Jev output, any sealed benchmark item, or any model response text. The data combines:
- RouterBench (MMLU, HellaSwag, GSM8K, WinoGrande and MBPP subsets only): measured correctness and cost of 11 LLMs per prompt.
- EmbedLLM (Apache-2.0): measured correctness of 112 open models per question.
- RouteLLM gpt4_dataset (Apache-2.0; only UltraChat, Anthropic-HH, FLAN and TruthfulQA prompts): GPT-4-judged adequacy of Mixtral's answer.
- Chatbot Arena 55k (Apache-2.0): human preference between a weaker and a stronger model.
- Synthetic router turns (ours, CC-BY-4.0): coding agents, IDE chat, support copilots, batch jobs and more, each with the router's own context summary. DeepSeek-V3.2 and Qwen3.8-27B wrote them. Two model families labelled them blind (author, plus Gemma-4-31B or GLM-5.1 as critic). We keep a row only when both agree on the category; any other field where they disagree is masked.
- Measured router-style tasks (ours, CC-BY-4.0): procedurally generated requests (maths, dates, code output, logic, tables, strings, money) whose gold answer code computes. We ran each one on Mistral-Nemo-12B, Gemma-4-31B and DeepSeek-V3.2 and recorded correctness, tokens, cost and latency.
We split by a hash of the normalised question text into train 86 %, calib 4 % and test 10 %. The same question can never appear in two splits. The pool owner checked the frozen files against a sealed router benchmark pool (exact match, 8-gram containment and MiniLM cosine) and found zero overlaps. Every source and its licence is listed in DATA-LICENCES.md.
Training: 2 epochs, AdamW (encoder 3e-5, heads 1e-3), batch 32, max 512 tokens, bf16, on one RTX PRO 6000 (33 minutes). Both arms together (this model and a Qwen3-0.6B arm) cost about USD 3.40 of GPU time.
Limits (read before relying on it)
- In-distribution advantage. The held-out test shares its source distributions with the training data, while Laya and Jev are zero-shot. The router's own 70 held-out tasks are the cleanest out-of-distribution check. There Weiche gets category right 94.9 % of the time (Jev 79.7 %). Its difficulty ranking on those tasks is weak (rank correlation 0.19; Jev 0.30).
- The success heads predict abstract tiers. Small / mid / strong mean the panels above
(7B–13B, 30B–70B-class and GPT-4-class models). With a very different model line-up, use
the traits, or recalibrate
succon your own outcomes. - The route head is experimental. On the held-out test, exact cost arithmetic over the
success heads beat the learned route head (mean regret USD 0.035 vs 0.050 per decision). On
post-freeze tasks the route head was better (0.067 vs 0.083). The router computes cache and
latency economics exactly, so feeding it
succis the recommended use. - English-centric (ModernBERT). Other languages are untested.
int8is lossy for this architecture. Usefp16orint4.- Not a JevBench result and not evaluated on any sealed benchmark.
Licence
Apache-2.0, inherited from ModernBERT-large. The synthetic and measured data we created are CC-BY-4.0. Public datasets keep their own licences (see DATA-LICENCES.md).
Model tree for benchmarkheaven/weiche-395m
Base model
answerdotai/ModernBERT-large