GOVINDFROM's picture
Update file
c33a601 verified
|
Raw
History Blame Contribute Delete
3.28 kB
metadata
library_name: cbct-clinical-reasoner
tags:
  - medical-imaging
  - cbct
  - report-generation
  - toothfairy4
  - odin2026

cbct-clinical-reasoner

Finding predictor and calibrated report decoder for ODIN 2026 Task 1 (ToothFairy4): generating a maxillofacial surgical-planning report from a single 3D CBCT volume.

For each of ~989 clinician-written statements the model predicts the probability that it applies to this scan; per-statement thresholds fitted on out-of-fold predictions then decide which to emit. Selecting from observed clinician phrasing means an entailment failure can only come from choosing the wrong finding, never from invented language.

Contents

Path Description
bundle/prototypes.json Sentence-prototype label space built from the training corpus
bundle/decoder.json Per-prototype thresholds calibrated on out-of-fold predictions
bundle/checkpoints/ Per-fold encoder checkpoints, ensembled at inference
bundle/config.json Preprocessing and training configuration
bundle/fallback_report.txt Prior-only report used if inference fails
plots/ Diagnostic figures for every pipeline stage
results.json, RESULTS.md Metrics with provenance

Development metrics

Out-of-fold over 622 public training cases.

metric value
final (0.8 clinical + 0.2 captioning) 0.3544
clinical F1 (offline surrogate) 0.3861
logical precision 0.4910
logical recall 0.3181
BLEU-4 (grader-exact) 0.1493
METEOR (grader-exact) 0.3064

Ablation

variant final clinical BLEU-4 METEOR
corpus prior (no imaging) 0.3575 0.3937 0.1170 0.3082
linear, 122 features 0.3544 0.3861 0.1493 0.3064
fine-tuned encoder (29M params) 0.3403 0.3752 0.1331 0.2684

Mean out-of-fold mAP across folds: 0.0738

How to read these numbers

  • BLEU-4 and METEOR are exact. They reimplement the grader's own local implementations and match NLTK to machine precision, so they compare directly with the public leaderboard.
  • The clinical score is a surrogate. RadFact carries 80% of the challenge ranking, but computing it needs an LLM judge. The figure above comes from an offline lexical entailment model built to rank decoder variants cheaply. It is not the challenge metric — obtain that with cbct-reasoner evaluate --radfact-lite.
  • These are in-domain. Stratified folds over the public release, not the hidden 50-case external-centre test set. Leave-one-centre-out (--strategy center) is the honest external estimate and reads lower.

Intended use and limitations

Research artifact for a benchmark. Not a medical device. Generated text is a draft for review by a qualified clinician and must not be used for patient care. It has not been validated for any clinical purpose.

Data provenance

Derived from the access-controlled ToothFairy4 release. prototypes.json contains sentences taken verbatim from clinical reports, so this repository is private by default and must not be made public without checking the ToothFairy4 data-use agreement.