Duan
Duan is a fine-tuned Laya System-1 decision engine for emergency and clinical decision tasks. It answers a question about a patient state in a single forward pass — no autoregressive decoding, no chain-of-thought — and returns a typed decision together with a calibrated probability distribution, a confidence estimate, and an abstention signal.
Architecture
| Backbone encoder | jhu-clsp/mmBERT-base (ModernBERT, 768 hidden, 322 M params) |
| Head | 2-layer Transformer encoder + typed decision scorer (Embedding(3, d) type bias + LayerNorm/Linear/GELU/Linear scorer) |
| Sequence budget | max_len = 1024, head_max_len = 256 |
| Task format | [CLS] <type> instructions [SEP] [MASK] opt₀ … [MASK] optₙ [SEP] state [SEP] |
Three question types, sharing one scorer — the only structural difference is a 3-row type embedding:
| type | criteria |
gold keys |
answer |
|---|---|---|---|
choice |
{label: description} |
same labels | argmax option + distribution |
score |
[level 0 desc, level 1 desc, …] |
"0" … "n-1" |
expected level + distribution |
noul |
{"false": …, "true": …} |
false / true |
P(true) |
Output
Decision · Probability · Confidence · Abstention — a full distribution over options, an expected level for score, P(true) for noul, plus a calibrated confidence and an action head that can abstain rather than guess.
Training
| Objective | RLCD — group sampling from perturbed logits → strictly proper scoring reward (log score + spherical, + RPS for score) → group-normalised advantage × policy gradient, plus soft-label cross-entropy |
| Steps | 12,992 optimizer steps (8 epochs) |
| Effective batch | 64 sequences/step |
| Data | 103,886 train items + 1,500 held-out for calibration (stratified) |
| Hardware | 1 × A800 80 GB, bf16, 5.2 h |
| Seeds | 20260925 (train) / 20260922 (calibration split) |
Training data (8 sources, 3 languages)
| Source | Task | Lang | Licence |
|---|---|---|---|
| NHAMCS 2018–2022 (CDC/NCHS) | ED triage (ESI 1–5) | en | MIT (US public data) |
| UCI Diabetes 130-US hospitals | readmission window (score, 3 levels) | en | CC BY 4.0 |
| MedQA — US / Mainland / Taiwan | medical MCQ | en / zh-Hans / zh-Hant | see upstream dataset |
| CARE-Bench | care escalation (4 levels) | en | CC BY-NC 4.0 |
| DDXPlus (English) | differential-diagnosis re-ranking | en | CC BY 4.0 |
| PubMedQA PQA-L | conclusion supported (noul) | en | MIT |
⚠️ No source's raw text is redistributed here — only the resulting weights. Several sources are not clinical outcomes: MedQA labels are exam answer keys, CARE-Bench labels are human-annotated consensus grades, PubMedQA labels are literature annotations, and DDXPlus is fully synthetic.
Evaluation (test split, n = 10,538)
| Task | n | acc | majority | macro-F1 | ECE | other |
|---|---|---|---|---|---|---|
ddx_diagnosis |
2000 | 0.9905 | 0.1660 | 0.9900 | 0.0071 | synthetic data |
ed_triage (choice, 5) |
2000 | 0.5600 | 0.4975 | 0.3490 | 0.1423 | QWK 0.4290 |
readmission_window (score, 3) |
2000 | 0.5810 | 0.5350 | 0.4178 | 0.0391 | QWK 0.2239, MAE 0.4875, ±1 level 93.2 % |
medqa_mcq |
4415 | 0.3647 | 0.2258 | 0.3659 | 0.4564 | en 0.2620 / zh-Hans 0.4505 / zh-Hant 0.3220 (acc) |
care_escalation |
79 | 0.6456 | 0.3418 | 0.6317 | 0.2951 | small n |
pubmedqa (noul) |
44 | 0.7273 | 0.6364 | 0.6966 | 0.2727 | AUROC 0.6708, PR-AUC 0.8135 |
Read these carefully. Majority-class baselines are given because every task is imbalanced; macro-F1 is the primary metric; cross-task macro-F1 is not comparable (different label spaces). ddx_diagnosis is high because DDXPlus is synthetic and its option set comes from the generator's own posterior — this is closed-set re-ranking, not open-set diagnosis.
Calibration
rl_agent_config.json ships a temperature_by_options table keyed by "<type>:<n_options>" and an additional temperature_by_qid table keyed by "<qid>|<type>:<n_options>".
The per-qid table exists because the coarse bucket key cannot separate tasks that share a bucket — MedQA (5 options), NHAMCS triage (5 options) and CARE-Bench (4 options) all fall into choice:3-5, while their uncalibrated ECE differs by an order of magnitude.
To use it, rescale probabilities offline — the model emits p = softmax(logits / t), so p_new ∝ p^(t_old / t_new):
def rescale(probs, t_from, t_to):
e = t_from / t_to
w = {k: max(v, 1e-12) ** e for k, v in probs.items()}
s = sum(w.values())
return {k: v / s for k, v in w.items()}
Values are clamped to [0.5, 5] on load; any out-of-range fitted value is also recorded under temperature_by_qid_raw.
Usage
import laya
agent = laya.Agent("wipen/Duan")
res = agent.predict(
state={"age": 67, "chief_complaint": "chest pain", "heart_rate": 118, "spo2": 91, ...},
questions={"acuity": {
"type": "choice",
"instructions": "What ESI acuity level should this visit be triaged to?",
"criteria": {"1": "requires immediate life-saving intervention", ...},
}},
)
print(res["answers"]["acuity"]["choice"], res["answers"]["acuity"]["probability"])
Licence
Weights released under CC BY-NC 4.0, because one training source (CARE-Bench) is licensed CC BY-NC 4.0 and another (DDXPlus) requires attribution.
The base model (convaiinnovations/laya-multilingual) and the training code are Apache-2.0.
Commercial use is not permitted under this licence. If you need commercial rights, the CARE-Bench portion must be removed and the model retrained, or separate permission obtained from that dataset's authors.
Citation
@misc{duan2026,
title = {Duan: A Healthcare System 1 Decision Model},
author = {Han, Weipeng},
year = {2026},
howpublished = {\url{https://github.com/wipen/Duan}},
note = {Work in progress}
}
If you use the underlying engine, please also cite Laya:
@misc{laya,
title = {Laya: fast, non-autoregressive System 1 decision engine},
author = {{ConvAI Innovations}},
howpublished = {\url{https://github.com/convaiinnovations/laya}},
note = {Apache-2.0}
}
Model tree for wipen/Duan
Base model
convaiinnovations/laya-multilingual