# Gợi ý vị trí chèn sơ đồ trong Methodology > Danh sách các chỗ nên thêm hình (ngoài Figure pipeline tổng quan ở §Overview đã có). Mỗi mục ghi rõ: **anchor** (chèn ở đâu trong `methodology (1).tex`), **nội dung sơ đồ**, **caption gợi ý**, **độ ưu tiên**. Bạn xem rồi chọn cái nào làm. > > Ký hiệu ưu tiên: ⭐⭐⭐ rất nên có (paper/thesis nào cũng vẽ) · ⭐⭐ nên có · ⭐ tuỳ chọn. --- ## A. Phần Data Preparation (§sec:data) ### A1. Sơ đồ data pipeline 3 phase ⭐⭐⭐ - **Anchor:** ngay sau đoạn mở đầu §Data Preparation (đoạn "...operates in three logical phases: selection / image transfer / resize and re-shard...", ~dòng 244–255), trước §Sources. - **Nội dung:** flow ngang 3 phase → output: `[SSD: CSV metadata] → Selection → [Cloud notebook] → Image transfer (PhysioNet) → [GPU host] → Resize + tar shard → Unified JSON`. Đánh dấu nơi chạy (SSD/cloud/GPU) và artefact trung gian. - **Caption:** "Three-phase offline data pipeline. Heavy linear-cost work (DICOM join, parsing, resize) runs once; the training loop only consumes the unified JSON + resized image tree." - *Tương đương Figure 2.3 của report mẫu — gần như bắt buộc.* ### A2. Funnel lọc 4 tầng (selection) ⭐⭐⭐ - **Anchor:** đầu §Selection Pipeline (sec:selection, ~dòng 283–288), hoặc ngay sau paragraph (d) Stratified sampling. - **Nội dung:** sơ đồ phễu (funnel) thu hẹp dần: `~227k studies → (a) frontal-only → (b) report parse (cần cả Findings+Impression) → (c) length-outlier removal → (d) stratified patient-disjoint sampling → 50k (40k/5k/5k)`. Ghi số study còn lại sau mỗi tầng nếu có. - **Caption:** "Four-stage filter chain reducing the full MIMIC-CXR corpus to the 50,000-study working subset." - *Rất "ăn điểm" và dễ hiểu; nên có số liệu thật ở mỗi tầng.* ### A3. Sơ đồ rẽ nhánh report_mode / image_mode → samples ⭐⭐ - **Anchor:** sau bảng `tab:reportmodes` (~dòng 418), trong §Unified Instruction JSON. - **Nội dung:** 1 study → (image_mode: frontal_only_split → 1 ảnh) → (report_mode: split_cascade → 2 sample: `findings` từ ảnh; `impression` từ ảnh + GT findings). Minh hoạ cascade nối findings→impression. - **Caption:** "How one study is serialised into training samples under `report_mode=split_cascade` and `image_mode=frontal_only_split`." --- ## B. Phần Model Architecture (§sec:arch) ### B1. RAD-DINO patchification ⭐⭐ - **Anchor:** trong §Image Encoder (sec:encoder), sau câu "...emits 1369 patch tokens of dimension 768... class token is discarded..." (~dòng 453–456). - **Nội dung:** `518×518 image → ViT-B/14 chia patch 14×14 → grid 37×37 = 1369 tokens (768-d) [+ CLS bị loại]`. Có thể thêm overlay heatmap minh hoạ patch. - **Caption:** "RAD-DINO patchification: a 518×518 chest X-ray becomes 1369 patch tokens (768-d); the [CLS] token is dropped before projection." ### B2. Chi tiết module MLP Projection ⭐⭐⭐ - **Anchor:** trong §MLP Projection (sec:projection), ngay sau khối align equations (eq:tap, ~dòng 501) hoặc cuối tiểu mục. - **Nội dung:** `Patch features P (1369×768) + 32 learnable query Q0 → CrossAttn 8-head → H0 (32×768) → Linear W1 + GELU + Dropout → H1 (32×1024) [TAP cho ITC head] → Linear W2 → V (32×4096)`. Đánh dấu nhánh 1024-d rẽ sang ITC head. - **Caption:** "MLP projection module. A 32-query cross-attention pool compresses 1369 patch tokens; the 1024-d intermediate is the tap point routed to the ITC head (Stage 1) or onward to the LLM (Stage 2)." - *Module trung tâm của model — nên vẽ kỹ.* ### B3. CheXpert classifier → PNU string ⭐⭐ - **Anchor:** trong §CheXpert Classifier (sec:chex), sau khối verbatim ví dụ PNU (~dòng 537) hoặc cuối tiểu mục. - **Nội dung:** `RAD-DINO [CLS] embedding → MLP head → 14×3 logits (per-label argmax) → format_pnu() → PNU 3-section string → chèn vào prompt giữa visual tokens và instruction`. Phân biệt nhánh train (GT từ CSV) vs eval (classifier tự predict). - **Caption:** "Auxiliary CheXpert head: the global RAD-DINO embedding is mapped to a 14×3 (Positive/Negative/Uncertain) tensor and serialised into the PNU prompt block." ### B4. LoRA injection trên attention block ⭐⭐ - **Anchor:** trong §Language Model and PEFT (sec:llm), sau đoạn mô tả LoRA trên q/k/v/o (~dòng 584–593). - **Nội dung:** 1 transformer block: base weight W (4-bit NF4, frozen) + adapter `ΔW = B·A` (rank 16, trainable) trên q/k/v/o; FFN để nguyên. Có thể thêm công thức `h = Wx + (α/r)·BAx`. - **Caption:** "LoRA adaptation: rank-16 low-rank adapters on q/k/v/o of every Vicuna block; the 4-bit base weights stay frozen, the feed-forward sublayers are untouched." ### B5. Anatomy của prompt + `` expansion ⭐⭐⭐ - **Anchor:** trong §Prompt Assembly (sec:prompt), sau khối verbatim prompt skeleton (~dòng 606) và/hoặc sau đoạn mô tả expand 31 vị trí (~dòng 608–615). - **Nội dung:** hai phần: 1. **Prompt anatomy** — thanh ngang chia slot, tô màu theo task: `[SYSTEM] [] [PNU block] [task context: GT findings / empty] [instruction] [ASSISTANT: target]`. 2. **Image-token expansion** — `1 placeholder → thay bằng 32 visual tokens → attention mask & label tensor nở thêm 31 vị trí; vùng visual + prompt + pad gán -100`. - **Caption:** "Prompt assembly and image-token expansion. The single `` placeholder is replaced by 32 visual tokens; attention mask, position ids, and labels are expanded by 31, with all non-target positions masked to −100." - *Quan trọng cho phần dễ sai nhất của pipeline (loss masking).* --- ## C. Phần Training Strategy (§sec:training) ### C1. Sơ đồ curriculum 3 stage ⭐⭐⭐ - **Anchor:** đầu §Training Strategy, sau đoạn intro (~dòng 665–674), trước §Stage 0. - **Nội dung:** 3 khối nối tiếp, mỗi khối ghi module frozen/trainable + loss: - `Stage 0`: RAD-DINO frozen → MLP head (BCE, U-MultiClass). - `Stage 1 (ITC)`: encoder frozen, projection+ITC head trainable, **LLM không load**, loss InfoNCE. - `Stage 2`: encoder+classifier+base-LLM frozen, projection+LoRA trainable, loss causal CE. Mũi tên: projection weights Stage 1 → nạp vào Stage 2. - **Caption:** "Two-stage training curriculum (preceded by the Stage 0 classifier). Stage 1 aligns the projection contrastively without loading Vicuna; Stage 2 loads the full model and instruction-tunes projection + LoRA." ### C2. Sơ đồ contrastive Stage 1 (kiểu CLIP) ⭐⭐⭐ - **Anchor:** trong §Stage 1 (sec:stage1), sau equation InfoNCE (eq:itc, ~dòng 720). - **Nội dung:** hai nhánh: - Image: `image → RAD-DINO → projection → ITC head → mean-pool → 128-d L2-norm`. - Text (offline): `findings (fallback impression) → CXR-BERT → 128-d L2-norm → cache {study_id → tensor}`. - Ma trận tương đồng `B×B`, đường chéo = positive, InfoNCE đối xứng. - **Caption:** "Stage 1 image–text contrastive alignment. Text embeddings are precomputed offline with CXR-BERT; only the projection + ITC head receive gradients." ### C3. Sơ đồ loss masking (label tensor) ⭐⭐ - **Anchor:** trong §Loss Masking and Image-Token Accounting (~dòng 781–797). - **Nội dung:** một dải token tuyến tính tô màu: `[system −100][32 visual −100][PNU −100][instruction −100][pad −100]` vs `[assistant response: loss được tính]`. Minh hoạ chỉ vùng response đóng góp cross-entropy. - **Caption:** "Label masking. Cross-entropy is computed strictly on assistant-response tokens; system prompt, visual span, PNU block, instruction, and padding are all set to −100." --- ## D. Phần Evaluation (§sec:eval) — nếu giữ trong Methodology ### D1. Sơ đồ luồng evaluation ⭐ - **Anchor:** đầu §Evaluation Protocol hoặc trong §Tasks and Held-Out Data (~dòng 877–896). - **Nội dung:** `test image → [CheXpert PNU predicted] (+ GT findings nếu impression) → model greedy decode → hypothesis → {NLG metrics: BLEU/ROUGE/METEOR/BERTScore} + {Clinical: CheXbert F1} + {VQA: EM/token-F1/LLM-judge}`. - **Caption:** "Evaluation flow per task, from test image to the metric families reported." ### D2. Sơ đồ CheXbert F1 ⭐ - **Anchor:** trong §Clinical Correctness (sec:clinical-f1, ~dòng 938–961). - **Nội dung:** `generated report + reference report → CheXbert labeler → 2 vectors 14 nhãn → binarise (−1,0→neg) → macro-F1/P/R`. - **Caption:** "Clinical F1: both generated and reference reports are labeled by CheXbert; the 14-label vectors are compared after binarisation." --- ## Tổng hợp ưu tiên | # | Sơ đồ | Mục | Ưu tiên | |---|---|---|---| | A1 | Data pipeline 3 phase | §Data Prep | ⭐⭐⭐ | | A2 | Funnel lọc 4 tầng | §Selection | ⭐⭐⭐ | | B2 | MLP Projection chi tiết | §Projection | ⭐⭐⭐ | | B5 | Prompt anatomy + `` expansion | §Prompt | ⭐⭐⭐ | | C1 | Curriculum 3 stage | §Training | ⭐⭐⭐ | | C2 | Contrastive Stage 1 (CLIP-style) | §Stage 1 | ⭐⭐⭐ | | A3 | report_mode/image_mode → samples | §JSON | ⭐⭐ | | B1 | RAD-DINO patchification | §Encoder | ⭐⭐ | | B3 | CheXpert → PNU | §CheXpert | ⭐⭐ | | B4 | LoRA injection | §LLM | ⭐⭐ | | C3 | Loss masking | §Masking | ⭐⭐ | | D1 | Luồng evaluation | §Eval | ⭐ | | D2 | CheXbert F1 | §Clinical F1 | ⭐ | > Nếu muốn, tôi có thể chèn sẵn các marker `% [INSERT FIGURE Xn: ...]` vào đúng vị trí trong `methodology (1).tex` (chỉ là comment, không ảnh hưởng biên dịch) để bạn thấy ngay chỗ nào cần hình khi mở file.