laya-burger
An unofficial Traditional Chinese (Taiwan) fine-tune of the Laya decision model. Give it a state (text or
JSON) and typed questions (choice, score, yes/no noul); it returns a calibrated probability for every option
in one forward pass. It never generates text. It supports a 2,048-token input budget (the original ships with
1,024).
- Base:
convaiinnovations/laya, subfoldermultilingual(mmBERT-base encoder, 322M parameters), revision5e7b2b1b. - Languages: Traditional Chinese as written in Taiwan, and English. Simplified Chinese is not supported.
- Use for: request routing, customer-service triage, intent, sentiment and ordinal ratings, workplace message triage, structured document review.
- Code, evaluation harness and full results: https://github.com/xamjiang/laya-burger
The name is a nod to Taiwanese breakfast shops. Not affiliated with or endorsed by the Laya authors, Google, NARLabs / TAIDE, or the breakfast chain 拉亞漢堡 (Laya Burger).
Quickstart
pip install laya
import laya
agent = laya.Agent("xamjiang/laya-burger") # or a local folder; max_len 2048 is read from the config
state = {"message": "你好,我上週在網路上訂的冷氣到現在還沒送來,客服電話一直打不通,再不處理我要退貨了。"}
questions = {
"category": {"type": "choice", "instructions": "這則訊息屬於哪一類?",
"criteria": {"shipping": "出貨或配送問題", "billing": "付款或發票問題",
"product": "商品本身的問題", "other": "其他"}},
"urgent": {"type": "noul", "instructions": "這件事需要今天處理嗎?"},
"upset": {"type": "score", "instructions": "寄件人有多不滿?",
"criteria": ["平靜", "有點困擾", "明顯不滿", "非常生氣"]},
}
out = agent.system_one(state, questions)
a = out["answers"]
print("category:", a["category"]["choice"], a["category"]["probabilities"])
print("urgent P(true):", a["urgent"]["noul"])
print("upset (expected level 0-3):", a["upset"]["score"], a["upset"]["probabilities"])
Verified with a local folder on laya 0.3.11 and 0.3.26 (identical outputs); loading by Hub id was not tested
before upload. laya.Router accepts only its three official names: use laya.Agent(path_or_hub_id), or
Router(models={"multilingual": path}).

Results

Accuracy (%) with 95% bootstrap intervals; Δ in percentage points, paired; bold = interval excludes 0. The original gives identical results at 1,024 and 2,048 tokens on these sets. Majority is an oracle baseline. GPT intervals are in the evaluation write-up linked below.
| benchmark | n | random / majority | original @1024 | laya-burger @2048 | Δ [95% CI] | GPT-6-sol (reference) |
|---|---|---|---|---|---|---|
| MASSIVE intent zh-TW (20 options) | 2,974 | 5.0 / 7.0 | 45.3 [43.5, 47.1] | 53.8 [52.0, 55.6] | +8.4 [+6.9, +10.0] | 93.8 (n = 500) ‡ |
| Belebele zho_Hant (reading MCQ) | 900 | 25.0 / 27.9 | 34.6 [31.5, 37.6] | 34.0 [30.9, 37.1] | −0.6 [−3.9, +2.7] | 95.8 |
| SIB-200 zho_Hant (topic) | 204 | 14.3 / 25.0 | 75.5 [69.6, 81.4] | 74.0 [68.1, 79.9] | −1.5 [−6.9, +3.4] | 89.7 |
XNLI zh (OpenCC s2twp) |
5,010 | 33.3 / 33.3 | 72.5 [71.3, 73.7] | 70.2 [69.0, 71.5] | −2.3 [−3.3, −1.2] | 73.0 (n = 1,000) |
| TMMLU+ (knowledge QA; limitation) | 1,320 | 25.0 / 25.9 | 27.5 [25.2, 29.9] | 27.0 [24.7, 29.5] | −0.5 [−3.3, +2.3] | 91.4 |
| MASSIVE intent en | 2,974 | 5.0 / 7.0 | 57.6 [55.8, 59.3] | 65.0 [63.3, 66.7] | +7.5 [+6.1, +8.8] | 95.4 (n = 500) ‡ |
| XNLI en | 5,010 | 33.3 / 33.3 | 82.6 [81.6, 83.7] | 81.6 [80.5, 82.7] | −1.0 [−1.8, −0.2] | 84.8 (n = 1,000) |
| typed-decisions test (en) | 2,000 | 31.8 / 49.8 | 34.4 [32.2, 36.7] | 34.6 [32.5, 36.8] | +0.1 [−1.4, +1.7] | 71.0 |
| ECE, MASSIVE zh-TW (lower is better) | 2,974 | 0.358 | 0.081 | – |
‡ GPT-6-sol was run on the first 500 items of each MASSIVE test set. On those same 500 items the Laya models score
58.6 (original @1024) / 63.8 (laya-burger @2048) on zh-TW and 65.8 / 71.2 on en, so read the gap to GPT on the same
items (laya_same_ids in eval/results/summary/reference_gpt.json in the GitHub repository).
- XNLI zh and XNLI en are significant losses. XNLI en stays inside the pre-registered non-inferiority margin (lower bound −1.8 pp > −2 pp).
- Belebele and TMMLU+ are near chance for both models: reading-comprehension MCQ and knowledge QA are not what this model does. typed-decisions (gold = a teacher's three-sample average) is below its majority baseline for both models and is a non-regression check only.
- GPT-6-sol (
reasoning_effort: none, zero-shot prompt following OpenAI's guidance, run on subsets) is a generative LLM reference, not a peer; cost and latency are not compared. - Internal sets (not released) favour laya-burger, because all of them were used to choose the checkpoint. On a Claude-written general zh-TW set, also used for checkpoint selection, accuracy rises from 0.492 to 0.717 (+0.225 [+0.185, +0.267], n = 710); on real Kaohsiung 1999 requests, which are in-domain for laya-burger and were used for checkpoint selection, from 0.508 to 0.745 (+0.237 [+0.198, +0.279], n = 577). Full list with intervals in the evaluation write-up.
Latency (median, one request, RTX 4000 Ada, bf16): 29 ms for 1 question and 95 ms for 3 questions on a 2,048-token input; the same as the original model.
Long inputs: accepted, not reliably understood
laya-burger supports a 2,048-token input budget, but it does not reliably find relevant content inside long unrelated text. In a post-hoc test (#9b, designed after the pre-registered test #9 was uninformative), one customer message placed among 512–4,096 tokens of unrelated passages costs laya-burger 0.14–0.46 accuracy in every cell, even when the message fully survives truncation. With the message at the start, it drops more than the original (−0.22 to −0.30 vs −0.03 to −0.07). The budget helps only when the relevant text would otherwise be cut off: at the end of a 1,024-token document, 0.300 vs 0.067 for the original at its default budget. Keep states focused.

Limitations

- It makes plain mistakes. Above, a review praises the product and complains only about shipping; laya-burger gives 2 stars (0.72) and P = 0.82 to "mentions a quality problem", confusing "mentions quality" with "mentions a quality problem". The scenario was written before the model ran; the output is unedited.
- XNLI zh −2.3 pp and XNLI en −1.0 pp compared with the original.
- Near chance on reading-comprehension MCQ and knowledge QA.
- Long unrelated text in the state hurts accuracy (see above).
- All labels come from one teacher model (Gemma 4 26B A4B); most states are synthetic. Single training seed.
- Checkpoint chosen with the internal sets, a 1,000-item MASSIVE test subset and typed-decisions test.
- Taiwan Mandarin and English only; no Simplified Chinese. Not for final decisions about people in medical, legal or hiring settings.
Training data and procedure
- One fine-tuning run from Laya multilingual, 2 epochs, 136,683 items (zh-TW 95,377, en 41,306), 2.66 h on one RTX 4000 Ada 20 GB.
- States written by Gemma 4 26B A4B (teacher; Apache-2.0) and by TAIDE
Gemma-3-TAIDE-12b-Chat-2602(20,518 of 130,939 train-split examples, 15.7%), plus real Taiwanese text: Kaohsiung 1999 requests and join.gov.tw citizen proposals (Taiwan government open data) and the first user turns of twllm-data (model replies discarded). Personal data masked; Simplified Chinese dropped. English replay: the typed-decisions train split with the original Laya's predictions as targets (its states come from a model the dataset authors do not name). - Soft labels from the teacher (blind relabelling, disagreement filtering); soft cross-entropy plus a ranked-probability loss on score questions; one temperature per question type fitted after training.
- No outputs of Claude, GPT or Gemini were trained on. No evaluation set was in training data or generation prompts.
- Details: TRAINING.md.
Evaluation
The protocol was registered before any full run; deviations and the post-hoc test #9b are documented. Every evaluation
number on this card comes from a results file in the GitHub repository (item ids, gold labels and probabilities
only); training numbers come from docs/training_facts.json there.
Contamination check: no shared n-grams with MASSIVE, Belebele, SIB-200 or XNLI zh; three TMMLU+ questions appear in
trained twllm-data user turns (one is in the evaluated sample). See
EVALUATION.md.
License and attribution
license: other because the licensing is composite; see LICENSE, NOTICE and LICENSES/.
- The weights are a derivative of the TAIDE G-class models and are distributed under the (Gemma edition) TAIDE
Model License Agreement, 2025-06-30 (「(Gemma 版次)-TAIDE 模型授權條款」(2025.06.30 版); copy in
LICENSES/), with the Apache-2.0 notice of Laya and the MIT notice of mmBERT-base. The Gemma Terms of Use are inLICENSES/too. - Binding use restrictions: no military or unlawful use; do not present outputs as human-generated; follow the Gemma Prohibited Use Policy.
本模型為財團法人國家實驗研究院 TAIDE G 類模型(https://taide.tw)之衍生模型。 This model is a derivative of the TAIDE G-class models of the National Applied Research Laboratories (https://taide.tw).
Gemma is provided under and subject to the Gemma Terms of Use found at ai.google.dev/gemma/terms
This model is an independent, unofficial project. It is not made, endorsed or reviewed by Convai Innovations or the Laya authors, the mmBERT authors, Google, or the National Applied Research Laboratories / TAIDE.
Citation
@misc{laya_burger_2026,
title = {laya-burger: an unofficial Traditional Chinese (Taiwan) fine-tune of the Laya decision model},
author = {SamJiang},
year = {2026},
howpublished = {\url{https://huggingface.co/xamjiang/laya-burger}},
note = {Version 1.0.0}
}
Please also cite Laya, mmBERT (arXiv:2509.06888), and the benchmarks you use; BibTeX for all of them is in the GitHub README.
繁體中文
Laya 決策模型的非官方台灣繁體中文微調版。 給它一個 state(純文字或 JSON)和幾個帶型別的問題(單選 choice、
等級 score、是非題 noul),它只要一次前向傳遞,就會為每個選項回傳校準過的機率。它從不生成文字。輸入長度上限為
2,048 tokens(原版為 1,024)。
- 基礎模型:
convaiinnovations/laya,子資料夾multilingual(mmBERT-base 編碼器,322M 參數),revision5e7b2b1b。 - 語言:台灣使用的繁體中文,以及英文。不支援簡體中文。
- 適用於:陳情案件分派、客服分流、意圖、情緒與等級評分、職場訊息分流、結構化文件審查。
- 程式碼、評測工具與完整結果:https://github.com/xamjiang/laya-burger
名稱是向台灣早餐店致意。本模型與 Laya 作者、Google、國家實驗研究院(NARLabs)/TAIDE,以及台灣早餐連鎖品牌 「拉亞漢堡」(Laya Burger)都沒有任何關係,也未獲其背書。
快速開始
pip install laya
import laya
agent = laya.Agent("xamjiang/laya-burger") # or a local folder; max_len 2048 is read from the config
state = {"message": "你好,我上週在網路上訂的冷氣到現在還沒送來,客服電話一直打不通,再不處理我要退貨了。"}
questions = {
"category": {"type": "choice", "instructions": "這則訊息屬於哪一類?",
"criteria": {"shipping": "出貨或配送問題", "billing": "付款或發票問題",
"product": "商品本身的問題", "other": "其他"}},
"urgent": {"type": "noul", "instructions": "這件事需要今天處理嗎?"},
"upset": {"type": "score", "instructions": "寄件人有多不滿?",
"criteria": ["平靜", "有點困擾", "明顯不滿", "非常生氣"]},
}
out = agent.system_one(state, questions)
a = out["answers"]
print("category:", a["category"]["choice"], a["category"]["probabilities"])
print("urgent P(true):", a["urgent"]["noul"])
print("upset (expected level 0-3):", a["upset"]["score"], a["upset"]["probabilities"])
已用本機資料夾在 laya 0.3.11 與 0.3.26 上驗證(輸出完全相同);上傳前尚未測試以 Hub id 載入。laya.Router
只接受它的三個官方名稱:請使用 laya.Agent(path_or_hub_id),或 Router(models={"multilingual": path})。

結果

準確率(%)附 95% bootstrap 信賴區間;Δ 為成對差異,單位為百分點(pp);粗體表示區間不含 0。原版在這些 評測集上,用 1,024 和 2,048 tokens 的結果完全相同。多數類是 oracle 基準。GPT 的信賴區間見下方連結的評測說明。
| 評測集 | n | 隨機 / 多數類 | 原版 @1024 | laya-burger @2048 | Δ [95% CI] | GPT-6-sol(參考) |
|---|---|---|---|---|---|---|
| MASSIVE intent zh-TW(20 個選項) | 2,974 | 5.0 / 7.0 | 45.3 [43.5, 47.1] | 53.8 [52.0, 55.6] | +8.4 [+6.9, +10.0] | 93.8 (n = 500) ‡ |
| Belebele zho_Hant(閱讀測驗選擇題) | 900 | 25.0 / 27.9 | 34.6 [31.5, 37.6] | 34.0 [30.9, 37.1] | −0.6 [−3.9, +2.7] | 95.8 |
| SIB-200 zho_Hant(主題分類) | 204 | 14.3 / 25.0 | 75.5 [69.6, 81.4] | 74.0 [68.1, 79.9] | −1.5 [−6.9, +3.4] | 89.7 |
XNLI zh(OpenCC s2twp) |
5,010 | 33.3 / 33.3 | 72.5 [71.3, 73.7] | 70.2 [69.0, 71.5] | −2.3 [−3.3, −1.2] | 73.0 (n = 1,000) |
| TMMLU+(知識問答;列為限制) | 1,320 | 25.0 / 25.9 | 27.5 [25.2, 29.9] | 27.0 [24.7, 29.5] | −0.5 [−3.3, +2.3] | 91.4 |
| MASSIVE intent en | 2,974 | 5.0 / 7.0 | 57.6 [55.8, 59.3] | 65.0 [63.3, 66.7] | +7.5 [+6.1, +8.8] | 95.4 (n = 500) ‡ |
| XNLI en | 5,010 | 33.3 / 33.3 | 82.6 [81.6, 83.7] | 81.6 [80.5, 82.7] | −1.0 [−1.8, −0.2] | 84.8 (n = 1,000) |
| typed-decisions test(en) | 2,000 | 31.8 / 49.8 | 34.4 [32.2, 36.7] | 34.6 [32.5, 36.8] | +0.1 [−1.4, +1.7] | 71.0 |
| ECE,MASSIVE zh-TW(越低越好) | 2,974 | 0.358 | 0.081 | – |
‡ GPT-6-sol 只跑了每個 MASSIVE 測試集的前 500 題。在同樣這 500 題上,兩個 Laya 模型在 zh-TW 為 58.6(原版 @1024)/
63.8(laya-burger @2048),en 為 65.8 / 71.2;與 GPT 的差距應以同一批題目比較(GitHub repository 中 eval/results/summary/reference_gpt.json 的 laya_same_ids)。
- XNLI zh 和 XNLI en 都是顯著退步。XNLI en 仍在事先登錄的非劣性界限之內(下界 −1.8 pp > −2 pp)。
- Belebele 和 TMMLU+ 上兩個模型都接近亂猜:閱讀測驗選擇題和知識問答不是這個模型的用途。typed-decisions(正確 標籤 = 教師模型三次取樣的平均)上兩個模型都低於多數類基準,只作為不退步檢查。
- GPT-6-sol(
reasoning_effort: none,依 OpenAI 指引撰寫的 zero-shot prompt,只在子集上執行)是生成式 LLM 的 參考點,不是同級對手;成本與延遲不列入比較。 - 內部評測集(不公開)對 laya-burger 有利,因為全部都被用來挑選 checkpoint。在一個由 Claude 撰寫、同樣用於挑選 checkpoint 的一般台灣中文評測集上,準確率從 0.492 升到 0.717(+0.225 [+0.185, +0.267],n = 710);在真實的 高雄 1999 陳情上(對 laya-burger 而言屬於同領域,也用於挑選 checkpoint),從 0.508 升到 0.745 (+0.237 [+0.198, +0.279],n = 577)。含信賴區間的完整清單見 評測說明。
延遲(中位數,單一請求,RTX 4000 Ada,bf16):2,048 tokens 輸入下,1 個問題 29 ms、3 個問題 95 ms;與原版模型 相同。
長輸入:讀得進去,但不一定讀得懂
laya-burger 支援 2,048 tokens 的輸入長度上限,但無法穩定地在冗長、不相關的文字中找出相關內容。在一個事後
追加的測試中(#9b,在事先登錄的測試 #9 無法提供資訊之後才設計),把一則顧客留言放進 512–4,096 tokens 的不相關
段落之間,laya-burger 在每一格的準確率都下降 0.14–0.46,即使那則留言完整保留、沒被截斷也一樣。留言放在開頭時,
它掉得比原版多(−0.22 到 −0.30,原版為 −0.03 到 −0.07)。長度上限只有在相關文字原本會被截掉時才有幫助:放在
1,024 tokens 文件的結尾時為 0.300,原版在預設上限下為 0.067。state 請保持聚焦。

限制

- 它會犯很單純的錯。上圖的評論稱讚商品,只抱怨出貨;laya-burger 給 2 星(0.72),並對「提到品質問題」給出 P = 0.82,把「提到品質」和「提到品質問題」混為一談。情境在模型執行前就已寫好;輸出未經修改。
- 與原版相比,XNLI zh −2.3 pp、XNLI en −1.0 pp。
- 閱讀測驗選擇題與知識問答接近亂猜。
state中不相關的長文會拉低準確率(見上文)。- 所有標籤都來自單一教師模型(Gemma 4 26B A4B);大部分
state是合成的。只有單一訓練隨機種子。 - checkpoint 是用內部評測集、一個 1,000 題的 MASSIVE 測試子集以及 typed-decisions test 挑選的。
- 只支援台灣華語與英文;不支援簡體中文。不適合在醫療、法律或招募情境中做關於人的最終決定。
訓練資料與流程
- 從 Laya multilingual 微調一次,2 個 epoch,136,683 筆(zh-TW 95,377、en 41,306),在單張 RTX 4000 Ada 20 GB 上 花了 2.66 小時。
state由 Gemma 4 26B A4B(教師模型;Apache-2.0)與 **TAIDEGemma-3-TAIDE-12b-Chat-2602**(130,939 個 train split 樣本中的 20,518 個,15.7%)撰寫,另加真實台灣文字:高雄 1999 陳情與 join.gov.tw 民眾提案(台灣政府 開放資料),以及 twllm-data 的第一則使用者訊息(模型 回覆全部捨棄)。個人資料已遮蔽;含簡體字的文字已捨棄。英文回放: typed-decisions 的 train split,以原版 Laya 的 預測作為目標值(它的state來自資料集作者沒有說明的某個模型)。- 教師模型給出的軟標籤(盲測重新標註、篩掉不一致的答案);soft cross-entropy,score 題另加 ranked-probability loss;訓練後每種題型各擬合一個溫度。
- 沒有拿任何 Claude、GPT 或 Gemini 的輸出來訓練。沒有任何評測集出現在訓練資料或生成 prompt 中。
- 細節:TRAINING.zh-TW.md。
評測
評測協議在任何完整執行之前就已登錄;與協議的偏差以及事後追加的測試 #9b 都有記錄。這張卡片上的每個評測數字,都來自
GitHub repository 中的結果檔(只包含題目 id、正確標籤和機率);訓練相關數字來自其中的 docs/training_facts.json。污染檢查:與 MASSIVE、Belebele、SIB-200 或
XNLI zh 沒有共用的 n-gram;有三題 TMMLU+ 題目出現在參與訓練的 twllm-data 使用者訊息中(其中一題在受評樣本裡)。
見 EVALUATION.zh-TW.md。
授權與出處標示
授權是複合式的,所以標為 license: other;見 LICENSE、NOTICE 與 LICENSES/。
- 權重屬於 TAIDE G 類模型的衍生模型,依「(Gemma 版次)-TAIDE 模型授權條款」(2025.06.30 版)散布
(英文名稱:(Gemma edition) TAIDE Model License Agreement, 2025-06-30;副本在
LICENSES/),並附上 Laya 的 Apache-2.0 聲明與 mmBERT-base 的 MIT 聲明。Gemma 使用條款的副本也在LICENSES/。 - 具約束力的使用限制:不得用於軍事或違法用途;不得把輸出當成人類產生的內容來呈現;須遵守 Gemma 禁止使用政策。
本模型為財團法人國家實驗研究院 TAIDE G 類模型(https://taide.tw)之衍生模型。 This model is a derivative of the TAIDE G-class models of the National Applied Research Laboratories (https://taide.tw).
Gemma is provided under and subject to the Gemma Terms of Use found at ai.google.dev/gemma/terms
本模型是獨立的非官方專案,並非由 Convai Innovations 或 Laya 作者、mmBERT 作者、Google,或財團法人國家實驗研究院/ TAIDE 製作、背書或審查。
引用
@misc{laya_burger_2026,
title = {laya-burger: an unofficial Traditional Chinese (Taiwan) fine-tune of the Laya decision model},
author = {SamJiang},
year = {2026},
howpublished = {\url{https://huggingface.co/xamjiang/laya-burger}},
note = {Version 1.0.0}
}
也請一併引用 Laya、mmBERT(arXiv:2509.06888),以及你使用的評測集;所有相關的 BibTeX 都在 GitHub README 中。
- Downloads last month
- 8
Model tree for xamjiang/laya-burger
Base model
convaiinnovations/laya