GOVINDFROM's picture
Update file
c33a601 verified
|
Raw
History Blame Contribute Delete
3.28 kB
---
library_name: cbct-clinical-reasoner
tags:
- medical-imaging
- cbct
- report-generation
- toothfairy4
- odin2026
---
# cbct-clinical-reasoner
Finding predictor and calibrated report decoder for **ODIN 2026 Task 1
(ToothFairy4)**: generating a maxillofacial surgical-planning report from a
single 3D CBCT volume.
For each of ~989 clinician-written statements the
model predicts the probability that it applies to this scan; per-statement
thresholds fitted on out-of-fold predictions then decide which to emit. Selecting
from observed clinician phrasing means an entailment failure can only come from
choosing the wrong finding, never from invented language.
## Contents
| Path | Description |
|---|---|
| `bundle/prototypes.json` | Sentence-prototype label space built from the training corpus |
| `bundle/decoder.json` | Per-prototype thresholds calibrated on out-of-fold predictions |
| `bundle/checkpoints/` | Per-fold encoder checkpoints, ensembled at inference |
| `bundle/config.json` | Preprocessing and training configuration |
| `bundle/fallback_report.txt` | Prior-only report used if inference fails |
| `plots/` | Diagnostic figures for every pipeline stage |
| `results.json`, `RESULTS.md` | Metrics with provenance |
## Development metrics
Out-of-fold over 622 public training cases.
| metric | value |
|---|---:|
| **final** (0.8 clinical + 0.2 captioning) | 0.3544 |
| clinical F1 (offline surrogate) | 0.3861 |
| logical precision | 0.4910 |
| logical recall | 0.3181 |
| BLEU-4 (grader-exact) | 0.1493 |
| METEOR (grader-exact) | 0.3064 |
## Ablation
| variant | final | clinical | BLEU-4 | METEOR |
|---|---:|---:|---:|---:|
| corpus prior (no imaging) | 0.3575 | 0.3937 | 0.1170 | 0.3082 |
| linear, 122 features | 0.3544 | 0.3861 | 0.1493 | 0.3064 |
| fine-tuned encoder (29M params) | 0.3403 | 0.3752 | 0.1331 | 0.2684 |
Mean out-of-fold mAP across folds: **0.0738**
## How to read these numbers
* **BLEU-4 and METEOR are exact.** They reimplement the grader's own local
implementations and match NLTK to machine precision, so they compare directly
with the public leaderboard.
* **The clinical score is a surrogate.** RadFact carries 80% of the challenge
ranking, but computing it needs an LLM judge. The figure above comes from an
offline lexical entailment model built to *rank* decoder variants cheaply. It
is not the challenge metric — obtain that with
`cbct-reasoner evaluate --radfact-lite`.
* **These are in-domain.** Stratified folds over the public release, not the
hidden 50-case external-centre test set. Leave-one-centre-out
(`--strategy center`) is the honest external estimate and reads lower.
## Intended use and limitations
Research artifact for a benchmark. **Not a medical device.** Generated text is a
draft for review by a qualified clinician and must not be used for patient care.
It has not been validated for any clinical purpose.
## Data provenance
Derived from the access-controlled ToothFairy4 release. `prototypes.json`
contains sentences taken verbatim from clinical reports, so this repository is
**private by default** and must not be made public without checking the
ToothFairy4 data-use agreement.