File size: 3,435 Bytes
f4ce534
7d982e1
 
f4ce534
 
 
7d982e1
 
 
f4ce534
 
7d982e1
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
0d0db9d
 
 
 
7d982e1
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3d0fe35
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
---
title: gradeeye
emoji: πŸ“š
colorFrom: pink
colorTo: green
sdk: static
pinned: true
thumbnail: >-
  https://cdn-uploads.huggingface.co/production/uploads/68cef9e7195e8b94d3557b81/CAfWeW32yUYWygr9dMb4Y.png
---

# GradEye

Model checkpoints and code supporting our study on **per-threshold calibration 
under domain shift** in diabetic retinopathy (DR) grading. We use a CORN 
ordinal regression head across a leave-one-domain-out (LODO) protocol on four 
public fundus photograph datasets, and show that calibration β€” not just 
accuracy β€” degrades non-uniformly across ordinal thresholds under domain shift.

πŸ”— Paper: coming soon
πŸ”— Code: [github.com/yourname/gradeeye](https://github.com/yourname/gradeeye)
πŸ“„ License: see below (per-repo, not MIT β€” derived from restricted-use source data)

## What we publish

Pretrained CNN/ViT backbones (ConvNeXt-Tiny, DeiT-3 Small/384, MaxViT-Tiny/384) 
with a CORN ordinal regression head, trained under a strict LODO protocol 
across four DR fundus photograph domains, plus auxiliary-channel variants 
(learned segmentation masks, Sobel edge maps) and the U-Net segmentation 
models used to generate them.

## Data sources (LODO folds)

| Domain | Role |
|:--|:--|
| EyePACS | held-out fold / testing source |
| APTOS | held-out fold / testing source |
| Messidor-2 | held-out fold / testing source |
| DDR | held-out fold / testing source |

Each model is trained on 3 of the 4 domains and evaluated on the 4th, 
held-out domain β€” the model never sees that domain during training or 
model selection.

## Key findings

| # | Finding | Headline number |
|--:|:--|:--|
| 1 | Per-threshold calibration diverges structurally under domain shift | ECE range 0.05–0.29 across thresholds (6Γ— spread on Messidor-2) |
| 2 | Post-hoc temperature scaling fails under shift β€” and fails the *opposite* way expected | T saturates at upper bound (5.0); model is under-confident, not overconfident |
| 3 | Class balancing hurts cross-domain accuracy in this LODO setup | unbalanced beats balanced by +3.93% mean QWK |
| 4 | Mild NPDR (grade 1) is the dominant cross-study failure mode | F1 < 0.13 on every fold, every variant |
| 5 | Best cross-domain backbone (of 3 tested) | ConvNeXt-Tiny, mean QWK 0.686 Β± 0.156 |

Full per-fold, per-variant tables are in the paper and in each repo's model card.

## Repository structure

One repo per training configuration (backbone Γ— channel variant Γ— loss). 
See the [Collection] *(link once live)* for the full index. Each repo's 
model card documents architecture, LODO fold, training hyperparameters, 
and metrics for that specific checkpoint set.

## License

Model weights are derived from datasets with **non-commercial, 
research/academic-use restrictions** (EyePACS and APTOS Kaggle competition 
terms prohibit redistribution and commercial use). Accordingly, all 
checkpoints in this org are released under **`cc-by-nc-4.0`**, research use 
only. See each repo's card for source-dataset attribution. Raw source 
images are not redistributed β€” only trained model weights.

## Status

This is an active research project; some tables in the paper are still 
being finalized (confidence intervals, seed sensitivity, full ablation 
matrix). Findings above are from completed runs and are stable; exact 
numbers may be refined before camera-ready.

## Author

[Ahmed Farhanur Rashid] Β· [GitHub](https://github.com/ahmed-farhanur-rashid)