Instructions to use keppy/evalroute-lane-encoder with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use keppy/evalroute-lane-encoder with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="keppy/evalroute-lane-encoder")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("keppy/evalroute-lane-encoder") model = AutoModelForSequenceClassification.from_pretrained("keppy/evalroute-lane-encoder", device_map="auto") - Notebooks
- Google Colab
- Kaggle
evalroute lane encoder v1
A small text classifier that maps a task request (the sentence you type to a coding agent) to one of evalroute's nine lanes. It is the optional first layer of evalroute's classifier, between the keyword rules and the LLM fallback.
Status: opt-in. On 19 real held-out tasks it ties the rules+LLM classifier and beats keyword rules alone, but neither result is distinguishable at 0.95. The decision below was made by gonogo, not by us.
evalroute install-encoder keppy/evalroute-lane-encoder # or a local dir
evalroute route "write the launch blog post" # card says: classification: encoder β¦
evalroute install-encoder --remove # back to rules (+ LLM if you have one)
Needs pip install "evalroute[encoder]" (torch CPU + transformers). Cold start without the
artifact installed does not import torch.
Numbers
Eval set: 19 real tasks held out from the maintainer's own routing ledger and the tier-a tasksets β 12 routine-coding, 4 dl-ml-research-engineering, 3 alignment-reasoning. Never a paraphrase, never seed text. Lanes with fewer than 8 real rows have no eval rows.
| classifier | accuracy | note |
|---|---|---|
| keyword rules only | 31.6% | what a standalone install has today; 10 of 19 hit no keyword |
| rules + LLM fallback | 68.4% | qwen/qwen3-coder-flash, one call per task |
| encoder v1 | 63.2% [41.0%, 80.9%] | $0 per route, no network |
gonogo compare (McNemar, paired by task id):
encoder vs rules: +31.6 pts, p = 0.11 β not distinguishable at 0.95.
encoder vs rules+LLM: β5.3 pts, p = 1.0 β tie.
gonogo decide (target 0.9): ASSIST ONLY β no confident subset reaches 90%.
Per lane (eval): alignment-reasoning 3/3, dl-ml 3/4, routine-coding 6/12. Every
routine-coding miss went to hard-agentic-coding or prose β two lanes the model has only
ever seen as paraphrases. Calibration: temperature 1.01, calibration error 0.18 (above
0.15 β treat defer_below 0.71 as a cut point, not a probability; it keeps 58% of rows at
β₯ 90% accuracy).
Training data (what is and is not here)
| source | rows | in this repo |
|---|---|---|
| LLM paraphrases of each lane's seed text (training only) | 191 | train_public.jsonl |
| tier-a taskset prompts | 23 | train_public.jsonl |
lane seed text (match_hint, keywords) |
18 | train_public.jsonl |
maintainer's own ledger routes, pinned with --lane |
29 | withheld β task text never leaves a machine |
split.json lists train/eval ids (ledger ids are opaque route ids). eval_predictions.jsonl
has id, label, predicted, confidence β no text. gonogo.txt is the verdict as printed.
Five of nine lanes (hard-agentic-coding, long-doc-reading, web-research, math-first-principles, orchestration) had zero real rows. v1 knows them only from paraphrases.
What we learned after shipping v1 (read this before using it)
A learning curve on the same 19 eval rows, run the morning after:
| training set | eval | routine-coding |
|---|---|---|
| paraphrases + seed + taskset, 0 real rows | 26% | 0/12 |
| β¦ + 14 real rows | 68% | 8/12 |
| β¦ + 29 real rows (= v1) | 47β63% | 4β6/12 |
| seed + taskset + 29 real rows, no paraphrases | 74% | 11/12 |
Paraphrases hurt the lanes they outnumber (the model learns what a lane sounds like when a model describes it) and are the only thing holding up the lanes with no real rows. Real rows are worth ~10Γ a paraphrase. And the maintainer's own ledger does not fill the empty lanes β 67% of its labels are coding dispatches.
So the direction changed: train your own encoder on your own ledger first (below β free, local, 30 s) and treat this shared model as a day-one floor. The shared encoder's data is now a public, human-reviewed corpus β https://huggingface.co/datasets/keppy/evalroute-tasks β real tasks with asserted lanes, published on purpose, no paraphrases. v2 will be trained from it, with paraphrases capped to the real-row count and only for lanes that still have none. If you want to improve the shared model, contribute reviewed rows there, especially for the five empty lanes.
Train your own
evalroute export-cases --out cases.jsonl --tasksets <dir> --seed-text
python scripts/split_cases.py --in cases.jsonl --train train.jsonl --eval eval.jsonl --split split.json
python examples/evalroute_lane_encoder.py --train train.jsonl --eval eval.jsonl --out my-encoder # thomas
evalroute install-encoder my-encoder
v1 trained in 30 s on a laptop CPU. The thomas recipe runs anywhere torch runs; --backend modal exists for larger sets. Compare yours against the shipped one on your own held-out
routes with gonogo before trusting it.
Files
model.safetensors, config.json, tokenizer*.json β ModernBERT-small-v2 fine-tune, 9 labels.
label2id.json, temperature.json, metrics.json β thomas CONTRACT Β§4 artifact.
train_public.jsonl, split.json, eval_predictions.jsonl, gonogo.txt β the evidence.
- Downloads last month
- 27
Model tree for keppy/evalroute-lane-encoder
Base model
johnnyboycurtis/ModernBERT-small-v2