Weiche-395M: a small local decision model for LLM auto-routing

Weiche (German for a railway switch) reads one user turn and decides which kind of model should handle it. It is ModernBERT-large (395M) with typed decision heads, trained for the open-source auto-model-router. One forward pass returns:

Output Meaning
category 9-way choice: coding, agentic, math, knowledge, long_context, tool_use, design, summarisation, general
scalars[1] difficulty rubric difficulty, 0 = trivial … 1 = frontier (the router's 0–4 scale ÷ 4)
scalars[0] outcome difficulty share of a model panel expected to fail (learned from measured outcomes)
scalars[2] stakes cost of an unnoticed wrong answer, 0–1
noul P(needs_tools), P(needs_vision), P(needs_long_context), P(follow_up)
succ P(success) of a small, a mid and a strong model tier, learned from measured outcomes
route experimental: small / mid / strong given 14 numeric router-state features (cache warm/cold, latency budget, value of a correct answer, tier prices and latencies)

It runs on CPU through ONNX Runtime (no PyTorch needed) and in the browser through onnxruntime-web (WebGPU).

Results

All numbers come from held-out data that was never used for training. The eval scripts and raw result JSON files are in eval/ and in the GitHub repo. The baselines are the router's existing decision paths: its heuristic, the local Laya 421M classifier (backend: local) and the hosted Jev API (backend: hosted). Jev was called only for this evaluation. Its outputs were never used for training, calibration or data selection.

The "+ fitted calibration" rows give the trait-only classifiers a per-tier logistic calibration from their traits to P(success). We fitted it on a separate calibration split, which is more help than the router's fixed success curve gives them. Weiche's success heads need no calibration. Routing picks, for each request, argmax P(success_t) − λ·cost_t and sweeps λ, as the router's expected-cost policy does. APGR is the average share of the quality gap between always-small and always-strong that is recovered across 20 cost budgets.

Classifier Category acc. (held-out, n=1134) Macro-F1 Category acc. on the router's own 70 held-out tasks needs_tools acc. Difficulty rank corr.
Weiche-395M (this model) 88.5 % 0.872 94.9 % (59 tasks) 98.4 % 0.739
Weiche-0.6B arm (Qwen3-0.6B, not released) 88.5 % 0.873 98.3 % (59 tasks) 99.2 % 0.736
hosted Jev 1.13 (API) 85.8 % 0.843 79.7 % (59 tasks) 96.0 % 0.753
Laya 421M (router local backend) 9.3 % 0.112 33.9 % (59 tasks) 80.5 % 0.297

RouterBench held-out (n=1500; tiers Mistral-7B / Mixtral-8x7B / GPT-4-1106)

Decision path Avg. gap recovered (APGR) Quality at 10 % of strong-tier cost at 25 % Cost share to recover 80 % of the gap 95 % of the gap
Router heuristic (no classifier) 0.494 47.2 % 47.2 % 100.0 % 100.0 %
Laya 421M (router local backend) + fitted calibration 0.812 48.9 % 53.8 % 38.9 % 79.5 %
hosted Jev 1.13 (API) + fitted calibration 0.850 50.3 % 59.6 % 28.2 % 86.2 %
Weiche-395M traits + fitted calibration 0.834 48.7 % 58.2 % 41.0 % 88.1 %
Weiche-395M success heads 0.862 51.4 % 59.0 % 31.4 % 65.2 %
Weiche-0.6B arm success heads 0.864 51.8 % 59.0 % 30.4 % 68.7 %

Always-small quality 28.4 %, always-mid 47.2 %, always-strong 68.6 %. An oracle that knows every outcome matches always-strong quality at 17.1 % of its cost and tops out at 77.2 %.

Measured router-style tasks finished after the data freeze (n=1103; tiers Mistral-Nemo-12B / Gemma-4-31B / DeepSeek-V3.2)

Decision path Avg. gap recovered (APGR) Quality at 10 % of strong-tier cost at 25 % Cost share to recover 80 % of the gap 95 % of the gap
Router heuristic (no classifier) 0.459 42.5 % 42.5 % 60.8 % 60.8 %
Laya 421M (router local backend) + fitted calibration 0.747 48.2 % 64.1 % 47.5 % 56.2 %
hosted Jev 1.13 (API) + fitted calibration 0.734 48.9 % 56.4 % 50.5 % 56.9 %
Weiche-395M traits + fitted calibration 0.772 52.4 % 62.4 % 44.0 % 54.4 %
Weiche-395M success heads 0.803 52.9 % 72.4 % 44.5 % 52.7 %
Weiche-0.6B arm success heads 0.805 54.0 % 71.4 % 41.9 % 52.7 %

Always-small quality 42.5 %, always-mid 95.0 %, always-strong 94.0 %. An oracle that knows every outcome matches always-strong quality at 38.3 % of its cost and tops out at 98.6 %.

CPU latency per decision, 4 threads, same busy host p50 p95
Weiche-395M ONNX int4 288 ms 393 ms
Weiche-395M ONNX int8 (lossy) 243 ms 382 ms
Weiche-395M ONNX fp16 386 ms 545 ms
Laya 421M (PyTorch, 7 questions) 11061 ms 12308 ms

The difficulty rank correlation compares against the averaged rubric labels of two model families. On that measure Weiche and hosted Jev are close (0.739 vs 0.753).

Use it in auto-model-router

policy:
  classifier: {backend: local-route-head, model: benchmarkheaven/weiche-395m, variant: fp16, threads: 2}

variant is fp16 (default, matches fp32 within 1e-3), int4 (423 MB, 96 % category agreement with fp32) or int8 (lossy for this architecture, 92 % agreement; not recommended). The first request downloads the chosen ONNX file once.

Use it directly (ONNX Runtime, Python)

import numpy as np, onnxruntime as ort
from huggingface_hub import snapshot_download
from tokenizers import Tokenizer

d = snapshot_download("benchmarkheaven/weiche-395m", allow_patterns=["model_fp16.onnx*", "tokenizer.json", "routerhead.py"])
tok = Tokenizer.from_file(f"{d}/tokenizer.json"); tok.enable_truncation(512)
sess = ort.InferenceSession(f"{d}/model_fp16.onnx", providers=["CPUExecutionProvider"])

def render(request, context=""):   # the training-time input format (see routerhead.py)
    if len(request) > 2400:
        request = request[:1400] + "\n[...]\n" + request[-1000:]
    return f"Context: {context.strip()[:600] or '(new conversation)'}\nRequest: {request}"

ids = np.array([tok.encode(render("Fix the flaky test in tests/test_api.py and run the suite",
                                  "3 messages, 2 from the user. Tools: Bash, Read, Edit.")).ids])
category, scalars, noul, succ, route = sess.run(None, {"input_ids": ids, "attention_mask": np.ones_like(ids),
                                                       "state": np.zeros((1, 14), np.float32)})

PyTorch: encoder/ holds the fine-tuned ModernBERT (HF format) and heads.safetensors holds the heads. routerhead.py defines RouterHead and render:

from transformers import AutoModel, AutoTokenizer
from safetensors.torch import load_file
from routerhead import RouterHead, probs, render
enc = AutoModel.from_pretrained("benchmarkheaven/weiche-395m", subfolder="encoder")
model = RouterHead(enc, 1024).eval()
model.load_state_dict(load_file("heads.safetensors"), strict=False)

Browser: model_int4.onnx ran under onnxruntime-web 1.30 with the WebGPU execution provider (eval/webgpu_smoke_int4.txt). Its outputs matched CPU ONNX Runtime within 1e-4. That check used Chrome's software WebGPU adapter, so it proves operator coverage, not speed. model_fp16.onnx needs a GPU with shader-f16; the software adapter has none, so the fp16 browser path is not yet verified on real hardware.

Training data

We did not use any TypeSafe/Jev output, any sealed benchmark item, or any model response text. The data combines:

  • RouterBench (MMLU, HellaSwag, GSM8K, WinoGrande and MBPP subsets only): measured correctness and cost of 11 LLMs per prompt.
  • EmbedLLM (Apache-2.0): measured correctness of 112 open models per question.
  • RouteLLM gpt4_dataset (Apache-2.0; only UltraChat, Anthropic-HH, FLAN and TruthfulQA prompts): GPT-4-judged adequacy of Mixtral's answer.
  • Chatbot Arena 55k (Apache-2.0): human preference between a weaker and a stronger model.
  • Synthetic router turns (ours, CC-BY-4.0): coding agents, IDE chat, support copilots, batch jobs and more, each with the router's own context summary. DeepSeek-V3.2 and Qwen3.8-27B wrote them. Two model families labelled them blind (author, plus Gemma-4-31B or GLM-5.1 as critic). We keep a row only when both agree on the category; any other field where they disagree is masked.
  • Measured router-style tasks (ours, CC-BY-4.0): procedurally generated requests (maths, dates, code output, logic, tables, strings, money) whose gold answer code computes. We ran each one on Mistral-Nemo-12B, Gemma-4-31B and DeepSeek-V3.2 and recorded correctness, tokens, cost and latency.

We split by a hash of the normalised question text into train 86 %, calib 4 % and test 10 %. The same question can never appear in two splits. The pool owner checked the frozen files against a sealed router benchmark pool (exact match, 8-gram containment and MiniLM cosine) and found zero overlaps. Every source and its licence is listed in DATA-LICENCES.md.

Training: 2 epochs, AdamW (encoder 3e-5, heads 1e-3), batch 32, max 512 tokens, bf16, on one RTX PRO 6000 (33 minutes). Both arms together (this model and a Qwen3-0.6B arm) cost about USD 3.40 of GPU time.

Limits (read before relying on it)

  • In-distribution advantage. The held-out test shares its source distributions with the training data, while Laya and Jev are zero-shot. The router's own 70 held-out tasks are the cleanest out-of-distribution check. There Weiche gets category right 94.9 % of the time (Jev 79.7 %). Its difficulty ranking on those tasks is weak (rank correlation 0.19; Jev 0.30).
  • The success heads predict abstract tiers. Small / mid / strong mean the panels above (7B–13B, 30B–70B-class and GPT-4-class models). With a very different model line-up, use the traits, or recalibrate succ on your own outcomes.
  • The route head is experimental. On the held-out test, exact cost arithmetic over the success heads beat the learned route head (mean regret USD 0.035 vs 0.050 per decision). On post-freeze tasks the route head was better (0.067 vs 0.083). The router computes cache and latency economics exactly, so feeding it succ is the recommended use.
  • English-centric (ModernBERT). Other languages are untested.
  • int8 is lossy for this architecture. Use fp16 or int4.
  • Not a JevBench result and not evaluated on any sealed benchmark.

Licence

Apache-2.0, inherited from ModernBERT-large. The synthetic and measured data we created are CC-BY-4.0. Public datasets keep their own licences (see DATA-LICENCES.md).

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for benchmarkheaven/weiche-395m

Quantized
(23)
this model