LLM Router — Stage 2 難度路由

判斷一則與工作相關的請求是單純還是複雜,據此決定要送往昂貴的雲端大模型(complex) 還是便宜的地端小模型(simple)。

  • 標籤:simple / complex
  • 基礎模型:xlm-roberta-base
  • 語言:繁體中文為主,中英混合

用途

在企業把請求送進昂貴的雲端大模型之前先判斷一次,避免兩種常見浪費: 用最貴的模型做分類、翻譯、單位換算這種小事;以及拿公司的 LLM 額度做與工作無關的事。

整個判斷在本地完成,不需要任何 API 金鑰,請求內容也不會送往第三方。

怎麼用

from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch

repo = "GOSHUNCLE/llm-router-complexity"
tok = AutoTokenizer.from_pretrained(repo)
mdl = AutoModelForSequenceClassification.from_pretrained(repo).eval()

text = "評估把回焊爐產能提升 20% 需要的資源與風險"
inputs = tok(text, return_tensors="pt", truncation=True, max_length=256)
with torch.no_grad():
    probs = torch.softmax(mdl(**inputs).logits, dim=-1)[0]
idx = int(probs.argmax())
print(mdl.config.id2label[idx], float(probs[idx]))

這個模型預設只處理已經被第一段判為 work_related 的請求。 直接拿無關的句子餵給它,結果沒有意義——請先過 相關性守門模型。

這是兩段式系統的一部分

階段 模型 標籤
第一段 相關性守門 GOSHUNCLE/llm-router-relevance-gate work_related / not_work_related
第二段 難度路由 GOSHUNCLE/llm-router-complexity simple / complex

完整的路由邏輯(含關鍵詞規則)、可直接試玩的 Demo、訓練腳本與資料集,都在 **GOSHUNCLE/liang-task-router**。

訓練資料

模板式合成資料,電子製造業情境、中英混合,不需要任何模型或金鑰即可重現 (training/generate_synthetic.py,固定 seed):

  • stage1.jsonl — 2,000 筆,work_related / not_work_related 各半
  • stage2.jsonl — 1,000 筆,simple / complex 各半

資料不含任何個人資料、公司名稱或真實使用者語料。

訓練設定

以 xlm-roberta-base 微調:4 epochs、learning rate 2e-5、batch size 16、max length 256, 資料切分 80 / 10 / 10(train / validation / test),依 Macro-F1 選最佳 checkpoint。

限制與注意事項

  • 領域侷限:訓練情境集中在電子製造業。其他產業的用語可能判不準。
  • 邊界句子會錯:「看起來簡單、實際需要推理」的句子最容易出錯。 實際遇過的例子是「分析 SMT 異常根因」被判成 simple。
  • 合成資料偏樂觀:模板生成的句子比真實使用者的輸入完整很多。
  • 不要單獨拿來做內容審查或合規判斷。它的用途是分流成本,不是把關風險。

授權

Apache License 2.0。基礎模型 xlm-roberta-base 為 MIT 授權。


English

Stage 2 of a two-stage LLM cost router: for work-related requests, decides whether the task is simple (route to a cheap/local model) or complex (route to an expensive cloud model). Fine-tuned from xlm-roberta-base on synthetic, template-generated data. Intended to run only on requests already passed by the Stage 1 relevance gate. No human-labeled benchmark exists yet, so no accuracy figures are claimed.

See the demo Space for the full routing logic (keyword rules + both stages), training scripts and dataset. Licensed under Apache-2.0; base model xlm-roberta-base is MIT.

Downloads last month
2
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for GOSHUNCLE/llm-router-complexity

Quantized
(38)
this model