loopback-kr commited on
Commit
7c1b186
·
verified ·
1 Parent(s): 983ded9

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +47 -11
README.md CHANGED
@@ -11,30 +11,58 @@ tags:
11
  - foundation-model
12
  - vl-bert
13
  - siglip2
14
- pipeline_tag: image-classification
 
 
15
  library_name: pytorch
16
  ---
17
 
18
  # Luna — Retinal Foundation Model with Hybrid Clinical & Anatomical Supervision
19
 
20
- **Luna** is a multimodal, multi-task retinal foundation model: a single-stream VL-BERT (SigLIP‑2 ViT‑Base vision encoder + PubMedBERT-initialized BERT‑Base text encoder) that reconstructs **masked clinical tokens** (disease label, age, sex, eye laterality) from 512×512 fundus images with a UNETR decoder jointly supervising vessel/optic-disc segmentation.
 
 
 
 
 
 
 
21
 
22
- > Paper: *Multimodal Foundation Model of High-Resolution Fundus Photos with Clinical Metadata via Predicting Demographics and Anatomic Structure* — Lim, H. et al.
23
  > Code: https://github.com/loopback-kr/Luna
24
 
25
  ## Model Details
26
 
27
  | Attribute | Value |
28
  |---|---|
29
- | **Architecture** | SigLIP‑2 ViT‑Base/16 vision encoder + BERT‑Base language encoder (single-stream VL-BERT fusion) + UNETR segmentation decoder |
30
- | **Parameters** | 93.5M (ViT‑Base backbone) |
31
  | **Input resolution** | 512 × 512 RGB (CLAHE-enhanced) |
32
- | **Vision encoder init.** | `google/siglip2-base-patch16-512` (ImageNet-pretrained) |
33
- | **Text encoder init.** | `NeuML/pubmedbert-base-embeddings`, extended with `[LABEL]`, `[AGE]`, `[SEX]`, `[DIR]` special tokens |
34
- | **Pre-training objective** | Masked clinical-token prediction (label/age/sex/direction) + vessel/optic-disc segmentation (Dice loss), *not* pixel-level MAE reconstruction |
35
  | **Pre-training data** | 361,519 CFPs (52K public + 309K private institutional images) |
36
  | **License** | MIT |
37
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
38
  ## Training Data
39
 
40
  **\*** = private institutional data, not available externally.
@@ -84,8 +112,16 @@ library_name: pytorch
84
 
85
  ```bibtex
86
  @article{lim_luna_2026,
87
- title = {Multimodal Foundation Model of High-Resolution Fundus Photos with Clinical Metadata via Predicting Demographics and Anatomic Structure},
88
- author = {Lim, Hyunseok and Kim, Junseok and Oh, Joonseo and Kim, Kanghyun and Lim, Jongsoo and Jeong, Jinhoon and Jeong, Hy and Kim, Yoonjeon and Kim, Namkug},
89
- journal = {(manuscript in preparation)},
 
 
90
  }
91
  ```
 
 
 
 
 
 
 
11
  - foundation-model
12
  - vl-bert
13
  - siglip2
14
+ - segmentation
15
+ - zero-shot-classification
16
+ pipeline_tag: image-feature-extraction
17
  library_name: pytorch
18
  ---
19
 
20
  # Luna — Retinal Foundation Model with Hybrid Clinical & Anatomical Supervision
21
 
22
+ **Luna** is a multimodal, multi-task retinal foundation model for high-resolution color fundus photographs (CFPs).
23
+ Instead of pixel reconstruction, it is pre-trained by **predicting masked clinical tokens** (disease label, age, sex,
24
+ eye laterality) from the image, while a **UNETR decoder jointly supervises vessel / optic-disc segmentation**.
25
+
26
+ Architecturally it is a single-stream VL-BERT: a **SigLIP-2 ViT-Base/16 (512px) vision encoder**, a **BERT-Base
27
+ masked-LM initialized from PubMedBERT**, and a **UNETR head** for anatomy.
28
+
29
+ > Paper: N/A
30
 
 
31
  > Code: https://github.com/loopback-kr/Luna
32
 
33
  ## Model Details
34
 
35
  | Attribute | Value |
36
  |---|---|
37
+ | **Architecture** | SigLIP-2 ViT-Base/16 vision encoder + BERT-Base masked-LM (single-stream VL-BERT fusion, 128 compressed vision tokens) + UNETR segmentation decoder |
38
+ | **Parameters** | 243M total; ≈93.5M in the SigLIP-2 ViT-Base vision encoder that downstream tasks reuse |
39
  | **Input resolution** | 512 × 512 RGB (CLAHE-enhanced) |
40
+ | **Vision encoder init.** | `google/siglip2-base-patch16-512` |
41
+ | **Text encoder init.** | `NeuML/pubmedbert-base-embeddings` tokenizer, extended with `[LABEL]`, `[AGE]`, `[SEX]`, `[DIR]` special tokens |
42
+ | **Pre-training objective** | Masked clinical-token prediction (label / age / sex / direction) + vessel & optic-disc Dice loss, *not* pixel-level MAE reconstruction |
43
  | **Pre-training data** | 361,519 CFPs (52K public + 309K private institutional images) |
44
  | **License** | MIT |
45
 
46
+ Prompt templates used during pre-training: `"A CFP image of {CLS}"`, `"age is {CLS} years old"`,
47
+ `"gender is {CLS}"`, `"this is {CLS} direction eye"`.
48
+
49
+ ## Files in this repository
50
+
51
+ This is the full `accelerate` training state at epoch 399:
52
+
53
+ | File | Contents |
54
+ |---|---|
55
+ | `model.safetensors` / `pytorch_model.bin` | Model weights (243,122,497 parameters, fp32) |
56
+ | `optimizer.bin`, `scheduler.bin`, `scaler.pt` | AdamW / LR-scheduler / GradScaler state, for exact resumption |
57
+ | `random_states_{0..3}.pkl` | RNG states of the four training processes |
58
+
59
+ ## Usage
60
+
61
+ Full training, fine-tuning and evaluation code lives in the
62
+ [GitHub repository](https://github.com/loopback-kr/Luna) — see its `README.md` for the *Upstream training*,
63
+ *Downstream training*, *Zero-Shot* and *Segmentation* sections. Checkpoint paths there expect
64
+ `pytorch_model.bin`; passing the containing directory instead resumes optimizer and scheduler state as well.
65
+
66
  ## Training Data
67
 
68
  **\*** = private institutional data, not available externally.
 
112
 
113
  ```bibtex
114
  @article{lim_luna_2026,
115
+ title = {Multimodal Foundation Model for High-Resolution Fundus Photographs Incorporating
116
+ Clinical Metadata via Prediction of Demographics and Anatomical Structures},
117
+ author = {Lim, Hyunseok and Kim, Junseok and Oh, Joonseo and Kim, Kanghyun and Lim, Jongsoo
118
+ and Jeong, Jinhoon and Jeong, Hy and Kim, Yoonjeon and Kim, Namkug},
119
+ year = {2026},
120
  }
121
  ```
122
+
123
+ ## License
124
+
125
+ Released under the MIT License — Copyright (c) 2026 Hyunseok Lim.
126
+ The third-party datasets listed above carry their own licenses and access terms, and the Asan Medical Center
127
+ institutional cohorts are not redistributable.