library_name: cbct-clinical-reasoner
tags:
- medical-imaging
- cbct
- report-generation
- toothfairy4
- odin2026
cbct-clinical-reasoner
Finding predictor and calibrated report decoder for ODIN 2026 Task 1 (ToothFairy4): generating a maxillofacial surgical-planning report from a single 3D CBCT volume.
For each of ~989 clinician-written statements the model predicts the probability that it applies to this scan; per-statement thresholds fitted on out-of-fold predictions then decide which to emit. Selecting from observed clinician phrasing means an entailment failure can only come from choosing the wrong finding, never from invented language.
Contents
| Path | Description |
|---|---|
bundle/prototypes.json |
Sentence-prototype label space built from the training corpus |
bundle/decoder.json |
Per-prototype thresholds calibrated on out-of-fold predictions |
bundle/checkpoints/ |
Per-fold encoder checkpoints, ensembled at inference |
bundle/config.json |
Preprocessing and training configuration |
bundle/fallback_report.txt |
Prior-only report used if inference fails |
plots/ |
Diagnostic figures for every pipeline stage |
results.json, RESULTS.md |
Metrics with provenance |
Development metrics
Out-of-fold over 622 public training cases.
| metric | value |
|---|---|
| final (0.8 clinical + 0.2 captioning) | 0.3544 |
| clinical F1 (offline surrogate) | 0.3861 |
| logical precision | 0.4910 |
| logical recall | 0.3181 |
| BLEU-4 (grader-exact) | 0.1493 |
| METEOR (grader-exact) | 0.3064 |
Ablation
| variant | final | clinical | BLEU-4 | METEOR |
|---|---|---|---|---|
| corpus prior (no imaging) | 0.3575 | 0.3937 | 0.1170 | 0.3082 |
| linear, 122 features | 0.3544 | 0.3861 | 0.1493 | 0.3064 |
| fine-tuned encoder (29M params) | 0.3403 | 0.3752 | 0.1331 | 0.2684 |
Mean out-of-fold mAP across folds: 0.0738
How to read these numbers
- BLEU-4 and METEOR are exact. They reimplement the grader's own local implementations and match NLTK to machine precision, so they compare directly with the public leaderboard.
- The clinical score is a surrogate. RadFact carries 80% of the challenge
ranking, but computing it needs an LLM judge. The figure above comes from an
offline lexical entailment model built to rank decoder variants cheaply. It
is not the challenge metric — obtain that with
cbct-reasoner evaluate --radfact-lite. - These are in-domain. Stratified folds over the public release, not the
hidden 50-case external-centre test set. Leave-one-centre-out
(
--strategy center) is the honest external estimate and reads lower.
Intended use and limitations
Research artifact for a benchmark. Not a medical device. Generated text is a draft for review by a qualified clinician and must not be used for patient care. It has not been validated for any clinical purpose.
Data provenance
Derived from the access-controlled ToothFairy4 release. prototypes.json
contains sentences taken verbatim from clinical reports, so this repository is
private by default and must not be made public without checking the
ToothFairy4 data-use agreement.