GOVINDFROM's picture
Update RESULTS.md
d4a79e4 verified
|
Raw
History Blame Contribute Delete
1.27 kB

Results

  • cases: 622
  • prototypes: 989 (tooth-aware=True, assign threshold=0.45)
  • out-of-fold mAP: 0.0738 (folds: 0.0883, 0.0667, 0.0694, 0.0805, 0.0640)

Out-of-fold score

metric value
final (0.8 clinical + 0.2 captioning) 0.3544
clinical F1 (surrogate) 0.3861
logical precision 0.4910
logical recall 0.3181
BLEU-4 (grader-exact) 0.1493
METEOR (grader-exact) 0.3064

Ablation

variant final clinical BLEU-4 METEOR
corpus prior (no imaging) 0.3575 0.3937 0.1170 0.3082
linear, 122 features 0.3544 0.3861 0.1493 0.3064
fine-tuned encoder (29M params) 0.3403 0.3752 0.1331 0.2684

Caveats

  • Clinical scores are the repository's offline RadFact surrogate unless 'evaluation_radfact_lite' is populated. The surrogate ranks decoder variants; it is not the challenge metric.
  • BLEU-4 and METEOR are exact reimplementations of the grader's local implementations and match NLTK to machine precision.
  • Scores are out-of-fold on the public training release, not the hidden 50-case test set, and the folds are stratified rather than leave-one-centre-out, so they measure in-domain performance.