cxr-vlm-code / docs /methodology_figure_placements.md
convitom
f
215d2b1
|
Raw
History Blame Contribute Delete
9.95 kB

Gợi ý vị trí chèn sơ đồ trong Methodology

Danh sách các chỗ nên thêm hình (ngoài Figure pipeline tổng quan ở §Overview đã có). Mỗi mục ghi rõ: anchor (chèn ở đâu trong methodology (1).tex), nội dung sơ đồ, caption gợi ý, độ ưu tiên. Bạn xem rồi chọn cái nào làm.

Ký hiệu ưu tiên: ⭐⭐⭐ rất nên có (paper/thesis nào cũng vẽ) · ⭐⭐ nên có · ⭐ tuỳ chọn.


A. Phần Data Preparation (§sec:data)

A1. Sơ đồ data pipeline 3 phase ⭐⭐⭐

  • Anchor: ngay sau đoạn mở đầu §Data Preparation (đoạn "...operates in three logical phases: selection / image transfer / resize and re-shard...", ~dòng 244–255), trước §Sources.
  • Nội dung: flow ngang 3 phase → output: [SSD: CSV metadata] → Selection → [Cloud notebook] → Image transfer (PhysioNet) → [GPU host] → Resize + tar shard → Unified JSON. Đánh dấu nơi chạy (SSD/cloud/GPU) và artefact trung gian.
  • Caption: "Three-phase offline data pipeline. Heavy linear-cost work (DICOM join, parsing, resize) runs once; the training loop only consumes the unified JSON + resized image tree."
  • Tương đương Figure 2.3 của report mẫu — gần như bắt buộc.

A2. Funnel lọc 4 tầng (selection) ⭐⭐⭐

  • Anchor: đầu §Selection Pipeline (sec:selection, ~dòng 283–288), hoặc ngay sau paragraph (d) Stratified sampling.
  • Nội dung: sơ đồ phễu (funnel) thu hẹp dần: ~227k studies → (a) frontal-only → (b) report parse (cần cả Findings+Impression) → (c) length-outlier removal → (d) stratified patient-disjoint sampling → 50k (40k/5k/5k). Ghi số study còn lại sau mỗi tầng nếu có.
  • Caption: "Four-stage filter chain reducing the full MIMIC-CXR corpus to the 50,000-study working subset."
  • Rất "ăn điểm" và dễ hiểu; nên có số liệu thật ở mỗi tầng.

A3. Sơ đồ rẽ nhánh report_mode / image_mode → samples ⭐⭐

  • Anchor: sau bảng tab:reportmodes (~dòng 418), trong §Unified Instruction JSON.
  • Nội dung: 1 study → (image_mode: frontal_only_split → 1 ảnh) → (report_mode: split_cascade → 2 sample: findings từ ảnh; impression từ ảnh + GT findings). Minh hoạ cascade nối findings→impression.
  • Caption: "How one study is serialised into training samples under report_mode=split_cascade and image_mode=frontal_only_split."

B. Phần Model Architecture (§sec:arch)

B1. RAD-DINO patchification ⭐⭐

  • Anchor: trong §Image Encoder (sec:encoder), sau câu "...emits 1369 patch tokens of dimension 768... class token is discarded..." (~dòng 453–456).
  • Nội dung: 518×518 image → ViT-B/14 chia patch 14×14 → grid 37×37 = 1369 tokens (768-d) [+ CLS bị loại]. Có thể thêm overlay heatmap minh hoạ patch.
  • Caption: "RAD-DINO patchification: a 518×518 chest X-ray becomes 1369 patch tokens (768-d); the [CLS] token is dropped before projection."

B2. Chi tiết module MLP Projection ⭐⭐⭐

  • Anchor: trong §MLP Projection (sec:projection), ngay sau khối align equations (eq:tap, ~dòng 501) hoặc cuối tiểu mục.
  • Nội dung: Patch features P (1369×768) + 32 learnable query Q0 → CrossAttn 8-head → H0 (32×768) → Linear W1 + GELU + Dropout → H1 (32×1024) [TAP cho ITC head] → Linear W2 → V (32×4096). Đánh dấu nhánh 1024-d rẽ sang ITC head.
  • Caption: "MLP projection module. A 32-query cross-attention pool compresses 1369 patch tokens; the 1024-d intermediate is the tap point routed to the ITC head (Stage 1) or onward to the LLM (Stage 2)."
  • Module trung tâm của model — nên vẽ kỹ.

B3. CheXpert classifier → PNU string ⭐⭐

  • Anchor: trong §CheXpert Classifier (sec:chex), sau khối verbatim ví dụ PNU (~dòng 537) hoặc cuối tiểu mục.
  • Nội dung: RAD-DINO [CLS] embedding → MLP head → 14×3 logits (per-label argmax) → format_pnu() → PNU 3-section string → chèn vào prompt giữa visual tokens và instruction. Phân biệt nhánh train (GT từ CSV) vs eval (classifier tự predict).
  • Caption: "Auxiliary CheXpert head: the global RAD-DINO embedding is mapped to a 14×3 (Positive/Negative/Uncertain) tensor and serialised into the PNU prompt block."

B4. LoRA injection trên attention block ⭐⭐

  • Anchor: trong §Language Model and PEFT (sec:llm), sau đoạn mô tả LoRA trên q/k/v/o (~dòng 584–593).
  • Nội dung: 1 transformer block: base weight W (4-bit NF4, frozen) + adapter ΔW = B·A (rank 16, trainable) trên q/k/v/o; FFN để nguyên. Có thể thêm công thức h = Wx + (α/r)·BAx.
  • Caption: "LoRA adaptation: rank-16 low-rank adapters on q/k/v/o of every Vicuna block; the 4-bit base weights stay frozen, the feed-forward sublayers are untouched."

B5. Anatomy của prompt + <image> expansion ⭐⭐⭐

  • Anchor: trong §Prompt Assembly (sec:prompt), sau khối verbatim prompt skeleton (dòng 606) và/hoặc sau đoạn mô tả expand 31 vị trí (dòng 608–615).
  • Nội dung: hai phần:
    1. Prompt anatomy — thanh ngang chia slot, tô màu theo task: [SYSTEM] [<image>] [PNU block] [task context: GT findings / empty] [instruction] [ASSISTANT: target].
    2. Image-token expansion1 <image> placeholder → thay bằng 32 visual tokens → attention mask & label tensor nở thêm 31 vị trí; vùng visual + prompt + pad gán -100.
  • Caption: "Prompt assembly and image-token expansion. The single <image> placeholder is replaced by 32 visual tokens; attention mask, position ids, and labels are expanded by 31, with all non-target positions masked to −100."
  • Quan trọng cho phần dễ sai nhất của pipeline (loss masking).

C. Phần Training Strategy (§sec:training)

C1. Sơ đồ curriculum 3 stage ⭐⭐⭐

  • Anchor: đầu §Training Strategy, sau đoạn intro (~dòng 665–674), trước §Stage 0.
  • Nội dung: 3 khối nối tiếp, mỗi khối ghi module frozen/trainable + loss:
    • Stage 0: RAD-DINO frozen → MLP head (BCE, U-MultiClass).
    • Stage 1 (ITC): encoder frozen, projection+ITC head trainable, LLM không load, loss InfoNCE.
    • Stage 2: encoder+classifier+base-LLM frozen, projection+LoRA trainable, loss causal CE. Mũi tên: projection weights Stage 1 → nạp vào Stage 2.
  • Caption: "Two-stage training curriculum (preceded by the Stage 0 classifier). Stage 1 aligns the projection contrastively without loading Vicuna; Stage 2 loads the full model and instruction-tunes projection + LoRA."

C2. Sơ đồ contrastive Stage 1 (kiểu CLIP) ⭐⭐⭐

  • Anchor: trong §Stage 1 (sec:stage1), sau equation InfoNCE (eq:itc, ~dòng 720).
  • Nội dung: hai nhánh:
    • Image: image → RAD-DINO → projection → ITC head → mean-pool → 128-d L2-norm.
    • Text (offline): findings (fallback impression) → CXR-BERT → 128-d L2-norm → cache {study_id → tensor}.
    • Ma trận tương đồng B×B, đường chéo = positive, InfoNCE đối xứng.
  • Caption: "Stage 1 image–text contrastive alignment. Text embeddings are precomputed offline with CXR-BERT; only the projection + ITC head receive gradients."

C3. Sơ đồ loss masking (label tensor) ⭐⭐

  • Anchor: trong §Loss Masking and Image-Token Accounting (~dòng 781–797).
  • Nội dung: một dải token tuyến tính tô màu: [system −100][32 visual −100][PNU −100][instruction −100][pad −100] vs [assistant response: loss được tính]. Minh hoạ chỉ vùng response đóng góp cross-entropy.
  • Caption: "Label masking. Cross-entropy is computed strictly on assistant-response tokens; system prompt, visual span, PNU block, instruction, and padding are all set to −100."

D. Phần Evaluation (§sec:eval) — nếu giữ trong Methodology

D1. Sơ đồ luồng evaluation ⭐

  • Anchor: đầu §Evaluation Protocol hoặc trong §Tasks and Held-Out Data (~dòng 877–896).
  • Nội dung: test image → [CheXpert PNU predicted] (+ GT findings nếu impression) → model greedy decode → hypothesis → {NLG metrics: BLEU/ROUGE/METEOR/BERTScore} + {Clinical: CheXbert F1} + {VQA: EM/token-F1/LLM-judge}.
  • Caption: "Evaluation flow per task, from test image to the metric families reported."

D2. Sơ đồ CheXbert F1 ⭐

  • Anchor: trong §Clinical Correctness (sec:clinical-f1, ~dòng 938–961).
  • Nội dung: generated report + reference report → CheXbert labeler → 2 vectors 14 nhãn → binarise (−1,0→neg) → macro-F1/P/R.
  • Caption: "Clinical F1: both generated and reference reports are labeled by CheXbert; the 14-label vectors are compared after binarisation."

Tổng hợp ưu tiên

# Sơ đồ Mục Ưu tiên
A1 Data pipeline 3 phase §Data Prep ⭐⭐⭐
A2 Funnel lọc 4 tầng §Selection ⭐⭐⭐
B2 MLP Projection chi tiết §Projection ⭐⭐⭐
B5 Prompt anatomy + <image> expansion §Prompt ⭐⭐⭐
C1 Curriculum 3 stage §Training ⭐⭐⭐
C2 Contrastive Stage 1 (CLIP-style) §Stage 1 ⭐⭐⭐
A3 report_mode/image_mode → samples §JSON ⭐⭐
B1 RAD-DINO patchification §Encoder ⭐⭐
B3 CheXpert → PNU §CheXpert ⭐⭐
B4 LoRA injection §LLM ⭐⭐
C3 Loss masking §Masking ⭐⭐
D1 Luồng evaluation §Eval
D2 CheXbert F1 §Clinical F1

Nếu muốn, tôi có thể chèn sẵn các marker % [INSERT FIGURE Xn: ...] vào đúng vị trí trong methodology (1).tex (chỉ là comment, không ảnh hưởng biên dịch) để bạn thấy ngay chỗ nào cần hình khi mở file.