evalroute lane encoder v1

A small text classifier that maps a task request (the sentence you type to a coding agent) to one of evalroute's nine lanes. It is the optional first layer of evalroute's classifier, between the keyword rules and the LLM fallback.

Status: opt-in. On 19 real held-out tasks it ties the rules+LLM classifier and beats keyword rules alone, but neither result is distinguishable at 0.95. The decision below was made by gonogo, not by us.

evalroute install-encoder keppy/evalroute-lane-encoder     # or a local dir
evalroute route "write the launch blog post"                # card says: classification: encoder …
evalroute install-encoder --remove                          # back to rules (+ LLM if you have one)

Needs pip install "evalroute[encoder]" (torch CPU + transformers). Cold start without the artifact installed does not import torch.

Numbers

Eval set: 19 real tasks held out from the maintainer's own routing ledger and the tier-a tasksets β€” 12 routine-coding, 4 dl-ml-research-engineering, 3 alignment-reasoning. Never a paraphrase, never seed text. Lanes with fewer than 8 real rows have no eval rows.

classifier accuracy note
keyword rules only 31.6% what a standalone install has today; 10 of 19 hit no keyword
rules + LLM fallback 68.4% qwen/qwen3-coder-flash, one call per task
encoder v1 63.2% [41.0%, 80.9%] $0 per route, no network

gonogo compare (McNemar, paired by task id): encoder vs rules: +31.6 pts, p = 0.11 β€” not distinguishable at 0.95. encoder vs rules+LLM: βˆ’5.3 pts, p = 1.0 β€” tie. gonogo decide (target 0.9): ASSIST ONLY β€” no confident subset reaches 90%.

Per lane (eval): alignment-reasoning 3/3, dl-ml 3/4, routine-coding 6/12. Every routine-coding miss went to hard-agentic-coding or prose β€” two lanes the model has only ever seen as paraphrases. Calibration: temperature 1.01, calibration error 0.18 (above 0.15 β€” treat defer_below 0.71 as a cut point, not a probability; it keeps 58% of rows at β‰₯ 90% accuracy).

Training data (what is and is not here)

source rows in this repo
LLM paraphrases of each lane's seed text (training only) 191 train_public.jsonl
tier-a taskset prompts 23 train_public.jsonl
lane seed text (match_hint, keywords) 18 train_public.jsonl
maintainer's own ledger routes, pinned with --lane 29 withheld β€” task text never leaves a machine

split.json lists train/eval ids (ledger ids are opaque route ids). eval_predictions.jsonl has id, label, predicted, confidence β€” no text. gonogo.txt is the verdict as printed.

Five of nine lanes (hard-agentic-coding, long-doc-reading, web-research, math-first-principles, orchestration) had zero real rows. v1 knows them only from paraphrases.

What we learned after shipping v1 (read this before using it)

A learning curve on the same 19 eval rows, run the morning after:

training set eval routine-coding
paraphrases + seed + taskset, 0 real rows 26% 0/12
… + 14 real rows 68% 8/12
… + 29 real rows (= v1) 47–63% 4–6/12
seed + taskset + 29 real rows, no paraphrases 74% 11/12

Paraphrases hurt the lanes they outnumber (the model learns what a lane sounds like when a model describes it) and are the only thing holding up the lanes with no real rows. Real rows are worth ~10Γ— a paraphrase. And the maintainer's own ledger does not fill the empty lanes β€” 67% of its labels are coding dispatches.

So the direction changed: train your own encoder on your own ledger first (below β€” free, local, 30 s) and treat this shared model as a day-one floor. The shared encoder's data is now a public, human-reviewed corpus β€” https://huggingface.co/datasets/keppy/evalroute-tasks β€” real tasks with asserted lanes, published on purpose, no paraphrases. v2 will be trained from it, with paraphrases capped to the real-row count and only for lanes that still have none. If you want to improve the shared model, contribute reviewed rows there, especially for the five empty lanes.

Train your own

evalroute export-cases --out cases.jsonl --tasksets <dir> --seed-text
python scripts/split_cases.py --in cases.jsonl --train train.jsonl --eval eval.jsonl --split split.json
python examples/evalroute_lane_encoder.py --train train.jsonl --eval eval.jsonl --out my-encoder   # thomas
evalroute install-encoder my-encoder

v1 trained in 30 s on a laptop CPU. The thomas recipe runs anywhere torch runs; --backend modal exists for larger sets. Compare yours against the shipped one on your own held-out routes with gonogo before trusting it.

Files

model.safetensors, config.json, tokenizer*.json β€” ModernBERT-small-v2 fine-tune, 9 labels. label2id.json, temperature.json, metrics.json β€” thomas CONTRACT Β§4 artifact. train_public.jsonl, split.json, eval_predictions.jsonl, gonogo.txt β€” the evidence.

Downloads last month
27
Safetensors
Model size
37.2M params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for keppy/evalroute-lane-encoder

Finetuned
(1)
this model

Datasets used to train keppy/evalroute-lane-encoder