cxr-vlm-code / docs /methodology_figure_placements.md
convitom
f
215d2b1
|
Raw
History Blame Contribute Delete
9.95 kB
# Gợi ý vị trí chèn sơ đồ trong Methodology
> Danh sách các chỗ nên thêm hình (ngoài Figure pipeline tổng quan ở §Overview đã có). Mỗi mục ghi rõ: **anchor** (chèn ở đâu trong `methodology (1).tex`), **nội dung sơ đồ**, **caption gợi ý**, **độ ưu tiên**. Bạn xem rồi chọn cái nào làm.
>
> Ký hiệu ưu tiên: ⭐⭐⭐ rất nên có (paper/thesis nào cũng vẽ) · ⭐⭐ nên có · ⭐ tuỳ chọn.
---
## A. Phần Data Preparation (§sec:data)
### A1. Sơ đồ data pipeline 3 phase ⭐⭐⭐
- **Anchor:** ngay sau đoạn mở đầu §Data Preparation (đoạn "...operates in three logical phases: selection / image transfer / resize and re-shard...", ~dòng 244–255), trước §Sources.
- **Nội dung:** flow ngang 3 phase → output: `[SSD: CSV metadata] → Selection → [Cloud notebook] → Image transfer (PhysioNet) → [GPU host] → Resize + tar shard → Unified JSON`. Đánh dấu nơi chạy (SSD/cloud/GPU) và artefact trung gian.
- **Caption:** "Three-phase offline data pipeline. Heavy linear-cost work (DICOM join, parsing, resize) runs once; the training loop only consumes the unified JSON + resized image tree."
- *Tương đương Figure 2.3 của report mẫu — gần như bắt buộc.*
### A2. Funnel lọc 4 tầng (selection) ⭐⭐⭐
- **Anchor:** đầu §Selection Pipeline (sec:selection, ~dòng 283–288), hoặc ngay sau paragraph (d) Stratified sampling.
- **Nội dung:** sơ đồ phễu (funnel) thu hẹp dần: `~227k studies → (a) frontal-only → (b) report parse (cần cả Findings+Impression) → (c) length-outlier removal → (d) stratified patient-disjoint sampling → 50k (40k/5k/5k)`. Ghi số study còn lại sau mỗi tầng nếu có.
- **Caption:** "Four-stage filter chain reducing the full MIMIC-CXR corpus to the 50,000-study working subset."
- *Rất "ăn điểm" và dễ hiểu; nên có số liệu thật ở mỗi tầng.*
### A3. Sơ đồ rẽ nhánh report_mode / image_mode → samples ⭐⭐
- **Anchor:** sau bảng `tab:reportmodes` (~dòng 418), trong §Unified Instruction JSON.
- **Nội dung:** 1 study → (image_mode: frontal_only_split → 1 ảnh) → (report_mode: split_cascade → 2 sample: `findings` từ ảnh; `impression` từ ảnh + GT findings). Minh hoạ cascade nối findings→impression.
- **Caption:** "How one study is serialised into training samples under `report_mode=split_cascade` and `image_mode=frontal_only_split`."
---
## B. Phần Model Architecture (§sec:arch)
### B1. RAD-DINO patchification ⭐⭐
- **Anchor:** trong §Image Encoder (sec:encoder), sau câu "...emits 1369 patch tokens of dimension 768... class token is discarded..." (~dòng 453–456).
- **Nội dung:** `518×518 image → ViT-B/14 chia patch 14×14 → grid 37×37 = 1369 tokens (768-d) [+ CLS bị loại]`. Có thể thêm overlay heatmap minh hoạ patch.
- **Caption:** "RAD-DINO patchification: a 518×518 chest X-ray becomes 1369 patch tokens (768-d); the [CLS] token is dropped before projection."
### B2. Chi tiết module MLP Projection ⭐⭐⭐
- **Anchor:** trong §MLP Projection (sec:projection), ngay sau khối align equations (eq:tap, ~dòng 501) hoặc cuối tiểu mục.
- **Nội dung:** `Patch features P (1369×768) + 32 learnable query Q0 → CrossAttn 8-head → H0 (32×768) → Linear W1 + GELU + Dropout → H1 (32×1024) [TAP cho ITC head] → Linear W2 → V (32×4096)`. Đánh dấu nhánh 1024-d rẽ sang ITC head.
- **Caption:** "MLP projection module. A 32-query cross-attention pool compresses 1369 patch tokens; the 1024-d intermediate is the tap point routed to the ITC head (Stage 1) or onward to the LLM (Stage 2)."
- *Module trung tâm của model — nên vẽ kỹ.*
### B3. CheXpert classifier → PNU string ⭐⭐
- **Anchor:** trong §CheXpert Classifier (sec:chex), sau khối verbatim ví dụ PNU (~dòng 537) hoặc cuối tiểu mục.
- **Nội dung:** `RAD-DINO [CLS] embedding → MLP head → 14×3 logits (per-label argmax) → format_pnu() → PNU 3-section string → chèn vào prompt giữa visual tokens và instruction`. Phân biệt nhánh train (GT từ CSV) vs eval (classifier tự predict).
- **Caption:** "Auxiliary CheXpert head: the global RAD-DINO embedding is mapped to a 14×3 (Positive/Negative/Uncertain) tensor and serialised into the PNU prompt block."
### B4. LoRA injection trên attention block ⭐⭐
- **Anchor:** trong §Language Model and PEFT (sec:llm), sau đoạn mô tả LoRA trên q/k/v/o (~dòng 584–593).
- **Nội dung:** 1 transformer block: base weight W (4-bit NF4, frozen) + adapter `ΔW = B·A` (rank 16, trainable) trên q/k/v/o; FFN để nguyên. Có thể thêm công thức `h = Wx + (α/r)·BAx`.
- **Caption:** "LoRA adaptation: rank-16 low-rank adapters on q/k/v/o of every Vicuna block; the 4-bit base weights stay frozen, the feed-forward sublayers are untouched."
### B5. Anatomy của prompt + `<image>` expansion ⭐⭐⭐
- **Anchor:** trong §Prompt Assembly (sec:prompt), sau khối verbatim prompt skeleton (~dòng 606) và/hoặc sau đoạn mô tả expand 31 vị trí (~dòng 608–615).
- **Nội dung:** hai phần:
1. **Prompt anatomy** — thanh ngang chia slot, tô màu theo task: `[SYSTEM] [<image>] [PNU block] [task context: GT findings / empty] [instruction] [ASSISTANT: target]`.
2. **Image-token expansion**`1 <image> placeholder → thay bằng 32 visual tokens → attention mask & label tensor nở thêm 31 vị trí; vùng visual + prompt + pad gán -100`.
- **Caption:** "Prompt assembly and image-token expansion. The single `<image>` placeholder is replaced by 32 visual tokens; attention mask, position ids, and labels are expanded by 31, with all non-target positions masked to −100."
- *Quan trọng cho phần dễ sai nhất của pipeline (loss masking).*
---
## C. Phần Training Strategy (§sec:training)
### C1. Sơ đồ curriculum 3 stage ⭐⭐⭐
- **Anchor:** đầu §Training Strategy, sau đoạn intro (~dòng 665–674), trước §Stage 0.
- **Nội dung:** 3 khối nối tiếp, mỗi khối ghi module frozen/trainable + loss:
- `Stage 0`: RAD-DINO frozen → MLP head (BCE, U-MultiClass).
- `Stage 1 (ITC)`: encoder frozen, projection+ITC head trainable, **LLM không load**, loss InfoNCE.
- `Stage 2`: encoder+classifier+base-LLM frozen, projection+LoRA trainable, loss causal CE.
Mũi tên: projection weights Stage 1 → nạp vào Stage 2.
- **Caption:** "Two-stage training curriculum (preceded by the Stage 0 classifier). Stage 1 aligns the projection contrastively without loading Vicuna; Stage 2 loads the full model and instruction-tunes projection + LoRA."
### C2. Sơ đồ contrastive Stage 1 (kiểu CLIP) ⭐⭐⭐
- **Anchor:** trong §Stage 1 (sec:stage1), sau equation InfoNCE (eq:itc, ~dòng 720).
- **Nội dung:** hai nhánh:
- Image: `image → RAD-DINO → projection → ITC head → mean-pool → 128-d L2-norm`.
- Text (offline): `findings (fallback impression) → CXR-BERT → 128-d L2-norm → cache {study_id → tensor}`.
- Ma trận tương đồng `B×B`, đường chéo = positive, InfoNCE đối xứng.
- **Caption:** "Stage 1 image–text contrastive alignment. Text embeddings are precomputed offline with CXR-BERT; only the projection + ITC head receive gradients."
### C3. Sơ đồ loss masking (label tensor) ⭐⭐
- **Anchor:** trong §Loss Masking and Image-Token Accounting (~dòng 781–797).
- **Nội dung:** một dải token tuyến tính tô màu: `[system −100][32 visual −100][PNU −100][instruction −100][pad −100]` vs `[assistant response: loss được tính]`. Minh hoạ chỉ vùng response đóng góp cross-entropy.
- **Caption:** "Label masking. Cross-entropy is computed strictly on assistant-response tokens; system prompt, visual span, PNU block, instruction, and padding are all set to −100."
---
## D. Phần Evaluation (§sec:eval) — nếu giữ trong Methodology
### D1. Sơ đồ luồng evaluation ⭐
- **Anchor:** đầu §Evaluation Protocol hoặc trong §Tasks and Held-Out Data (~dòng 877–896).
- **Nội dung:** `test image → [CheXpert PNU predicted] (+ GT findings nếu impression) → model greedy decode → hypothesis → {NLG metrics: BLEU/ROUGE/METEOR/BERTScore} + {Clinical: CheXbert F1} + {VQA: EM/token-F1/LLM-judge}`.
- **Caption:** "Evaluation flow per task, from test image to the metric families reported."
### D2. Sơ đồ CheXbert F1 ⭐
- **Anchor:** trong §Clinical Correctness (sec:clinical-f1, ~dòng 938–961).
- **Nội dung:** `generated report + reference report → CheXbert labeler → 2 vectors 14 nhãn → binarise (−1,0→neg) → macro-F1/P/R`.
- **Caption:** "Clinical F1: both generated and reference reports are labeled by CheXbert; the 14-label vectors are compared after binarisation."
---
## Tổng hợp ưu tiên
| # | Sơ đồ | Mục | Ưu tiên |
|---|---|---|---|
| A1 | Data pipeline 3 phase | §Data Prep | ⭐⭐⭐ |
| A2 | Funnel lọc 4 tầng | §Selection | ⭐⭐⭐ |
| B2 | MLP Projection chi tiết | §Projection | ⭐⭐⭐ |
| B5 | Prompt anatomy + `<image>` expansion | §Prompt | ⭐⭐⭐ |
| C1 | Curriculum 3 stage | §Training | ⭐⭐⭐ |
| C2 | Contrastive Stage 1 (CLIP-style) | §Stage 1 | ⭐⭐⭐ |
| A3 | report_mode/image_mode → samples | §JSON | ⭐⭐ |
| B1 | RAD-DINO patchification | §Encoder | ⭐⭐ |
| B3 | CheXpert → PNU | §CheXpert | ⭐⭐ |
| B4 | LoRA injection | §LLM | ⭐⭐ |
| C3 | Loss masking | §Masking | ⭐⭐ |
| D1 | Luồng evaluation | §Eval | ⭐ |
| D2 | CheXbert F1 | §Clinical F1 | ⭐ |
> Nếu muốn, tôi có thể chèn sẵn các marker `% [INSERT FIGURE Xn: ...]` vào đúng vị trí trong `methodology (1).tex` (chỉ là comment, không ảnh hưởng biên dịch) để bạn thấy ngay chỗ nào cần hình khi mở file.