YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Assignment 2 β€” Topics in AI for Healthcare (Spring 26)

Chest X-ray classification, lung segmentation, Grad-CAM analysis (COVID-19 Radiography Database) and polyp object detection (Kvasir-SEG).

Seed = 16 (last two digits of the roll number). The same seed appears in the code, in every split filename, and in the report.

Before submitting: rename this folder to your full roll number (mv RollNumber_16 2026XXXX16), set your roll number in report_notes.md, and zip the result.


1. Setup

Nothing to install if your environment already has the libraries below. Confirm with:

python run_tests.py

Its first section prints every library with its version and fails loudly on anything missing; the remaining sections run 66 self-contained correctness checks that need no dataset.

Libraries

Every entry is genuinely imported by the code β€” nothing here is speculative.

Package Used for Verified version
torch all models, training loops, autograd (Grad-CAM) 2.14.0
torchvision DenseNet-121, Faster R-CNN / RetinaNet, transforms.v2 0.29.0
torchxrayvision Model 1 β€” the frozen chest-X-ray-pretrained DenseNet 1.5.4
numpy numerics throughout 2.4.4
pandas split CSVs, per-image metric tables 3.0.2
scikit-learn stratified splits, confusion matrix, P/R/F1, AUROC 1.8.0
scipy HD95 / ASSD distance transforms, ANOVA, Wilcoxon 1.17.1
Pillow image and mask IO 12.1.1
opencv-python-headless affine warps for joint image/mask augmentation, lung-crop resize 4.13.0.92
matplotlib every plot and overlay figure 3.10.8
reportlab report.pdf assembly 4.4.10
tqdm progress bars β€” optional, the code degrades gracefully without it 4.70.0
kagglehub download_data.py β‰₯0.3

torchxrayvision additionally pulls in scikit-image, imageio and requests at import time; these are listed in requirements.txt so they are not a surprise.

If anything is missing: pip install -r requirements.txt, or pip install -r requirements-lock.txt for the exact pinned set. For a CUDA build, install torch / torchvision from https://pytorch.org/get-started/locally/ first.

Datasets

python download_data.py

This pulls both datasets through kagglehub and symlinks them into Dataset/:

Dataset Kaggle slug Size
COVID-19 Radiography Database tawsifurrahman/covid19-radiography-database ~800 MB
Kvasir-SEG debeshjha1/kvasirseg ~46 MB

You need Kaggle credentials β€” ~/.kaggle/kaggle.json, or the KAGGLE_USERNAME / KAGGLE_KEY environment variables, or an interactive login on first use. Use --copy if your filesystem does not support symlinks.

Nothing needs to be inside Dataset/: every script also falls back to the kagglehub cache on its own, and --data-root overrides both.


2. Run everything

bash run_all.sh            # full run
bash run_all.sh --quick    # tiny run, just to verify the pipeline end to end

Already-trained checkpoints are skipped, so the script is safe to re-run.

Or step by step

Part A β€” COVID-19 Radiography

cd classification_segmentation

# 2.1 data prep, QC audit, frozen splits, EDA plots
python pre_process.py

# 2.2 Task 1 β€” Model 1: frozen TorchXRayVision DenseNet
python train.py    --task classify --model xrv --epochs 25
python evaluate.py --task classify --ckpt ../Results/model_weights/cls_xrv_seed16.pt

# 2.2 Task 1 β€” Model 2: student DenseNet-121
python train.py    --task classify --model student --init-from imagenet --epochs 25
python evaluate.py --task classify --ckpt ../Results/model_weights/cls_student_imagenet_seed16.pt

# 2.2 the seeded ablation, run identically for both models
python train.py    --task classify --model xrv     --ablation lung_crop --epochs 25
python train.py    --task classify --model student --ablation lung_crop --epochs 25
python evaluate.py --task classify --ckpt ../Results/model_weights/cls_xrv_abl-lung_crop_seed16.pt
python evaluate.py --task classify --ckpt ../Results/model_weights/cls_student_imagenet_abl-lung_crop_seed16.pt
python evaluate.py --compare cls_xrv_seed16 cls_xrv_abl-lung_crop_seed16 \
                             cls_student_imagenet_seed16 cls_student_imagenet_abl-lung_crop_seed16

# 2.3 Task 2 β€” lung segmentation
python train.py    --task segment --epochs 40
python evaluate.py --task segment --ckpt ../Results/model_weights/seg_unet_seed16.pt

# 2.4 Task 3 β€” Grad-CAM vs lung mask
python gradcam_analysis.py \
    --ckpt-xrv     ../Results/model_weights/cls_xrv_seed16.pt \
    --ckpt-student ../Results/model_weights/cls_student_imagenet_seed16.pt \
    --n-images 12

Part B β€” Kvasir-SEG

cd object_detection

python pre_process.py                                   # 3.1 + 3.2

python train.py    --loss ce    --epochs 20             # 3.3 ablation arm 1
python evaluate.py --ckpt ../Results/model_weights/det_fasterrcnn_resnet50_fpn_ce_seed16.pt

python train.py    --loss focal --epochs 20             # 3.3 ablation arm 2
python evaluate.py --ckpt ../Results/model_weights/det_fasterrcnn_resnet50_fpn_focal-a0.25g2.0_seed16.pt

python evaluate.py --compare det_fasterrcnn_resnet50_fpn_ce_seed16 \
                             det_fasterrcnn_resnet50_fpn_focal-a0.25g2.0_seed16

Report

python make_report.py       # builds report.pdf from Results/ + report_notes.md

Correctness checks (fast, no dataset, run this before submitting)

python run_tests.py

Sections you have not written yet appear as yellow TODO boxes in the PDF, and the script prints the list of what is still missing.

Useful flags

Flag Where Meaning
--ablation {none,lung_crop,border_mask,gray_input,aug_strong,aug_weak} A train the single preprocessing change
--imbalance {none,class_weight,sampler,focal} A train class-imbalance handling (default class_weight)
--init-from {imagenet,scratch} A train Model 2 initialisation
--loss {ce,focal,weighted_ce} B train the ROI-head loss ablation
--score-thr B evaluate fixed confidence threshold (default: argmax-F1 from the sweep)
--data-root both pre_process explicit dataset path
--workers 0 any if your machine struggles with dataloader workers

Rough runtimes on one modern GPU: Part A classification ~15 min per run, segmentation ~35 min, Grad-CAM ~2 min, Part B ~30 min per arm. On CPU everything still runs but expect roughly 20–40Γ— longer, so use --quick to check the wiring first.


3. Layout

RollNumber_16/
β”œβ”€β”€ classification_segmentation/
β”‚   β”œβ”€β”€ common.py            config, datasets, transforms, metrics, plot helpers
β”‚   β”œβ”€β”€ pre_process.py       indexing, QC audit, frozen splits, EDA
β”‚   β”œβ”€β”€ model.py             XRVDenseNet (frozen) | StudentDenseNet | UNet | losses
β”‚   β”œβ”€β”€ train.py             classification + segmentation training
β”‚   β”œβ”€β”€ evaluate.py          test metrics, ROC, error analysis, overlays, comparisons
β”‚   └── gradcam_analysis.py  Grad-CAM for both models + lung-mask overlap metrics
β”œβ”€β”€ object_detection/
β”‚   β”œβ”€β”€ common.py            Kvasir dataset, augmentation, AP/mAP implementation
β”‚   β”œβ”€β”€ pre_process.py       annotation checks, frozen split, detection EDA
β”‚   β”œβ”€β”€ model.py             Faster R-CNN + the focal / weighted-CE loss swap
β”‚   β”œβ”€β”€ train.py             detector training
β”‚   └── evaluate.py          mAP suite, threshold sweep, failure analysis
β”œβ”€β”€ Dataset/{covid19,kvasir}/          datasets (or symlinks to the kagglehub cache)
β”œβ”€β”€ splits/
β”‚   β”œβ”€β”€ covid_classification_seed16.csv
β”‚   β”œβ”€β”€ covid_segmentation_seed16.csv
β”‚   └── kvasir_detection_seed16.csv
β”œβ”€β”€ Results/{metrics,plots,overlays,model_weights}/
β”œβ”€β”€ download_data.py   run_all.sh         run_tests.py
β”œβ”€β”€ requirements.txt   requirements-lock.txt
β”œβ”€β”€ report_notes.md    make_report.py     report.pdf
β”œβ”€β”€ .gitignore         .gitattributes
└── README.md

Checkpoints in Results/model_weights/ are large β€” upload them to Google Drive and put the link in the report, as the assignment instructs.


4. Design decisions you should be ready to defend

These are the choices a viva will probe. Each is implemented and commented in the code.

  • Splits are grouped by MD5 before stratifying. An exact duplicate of a training image can therefore never land in test. This dataset is an aggregation of several public sources and does contain repeats.
  • No patient identifier exists, so patient-level leakage cannot be fully excluded. Say so in the limitations section rather than claiming a clean split.
  • No horizontal flip on chest X-rays. Laterality is anatomically fixed (heart on the patient's left, gastric bubble left) and burned-in L/R markers are real signal. Mirroring produces impossible images. Kvasir frames do get h- and v-flips, because an endoscope view has no canonical orientation.
  • Masks are resized with NEAREST and binarised at 0.5; images use BILINEAR. Interpolating a label map invents intermediate values that do not exist.
  • Images are 299Γ—299, masks are 256Γ—256 in this dataset version. They share the same field of view, so both are resized to a common 224. The audit counts these mismatches explicitly rather than silently assuming alignment.
  • Model 1 is a genuine linear probe: every backbone parameter has requires_grad=False and the backbone is held in eval() mode so its BatchNorm running statistics stay frozen. It reports ~6.1 K trainable parameters out of ~7.0 M.
  • Class imbalance (Normal β‰ˆ 10 k vs Viral Pneumonia β‰ˆ 1.3 k) is handled with inverse-frequency weighted cross-entropy by default; --imbalance sampler and --imbalance focal are available as alternatives. Balanced accuracy and macro-F1 are reported alongside accuracy for exactly this reason.
  • Grad-CAM needs x.requires_grad_(True) for Model 1. With a fully frozen backbone and a non-grad input, autograd builds no graph through the conv stack and activation gradients would be None.
  • Boundary metrics (HD95, ASSD, boundary-F1 @2 px) are reported next to Dice. Dice is dominated by the interior of a large lung field and can look healthy while the pleural or diaphragm edge is systematically wrong.
  • The detection loss ablation changes exactly one thing. Only roi_heads.fastrcnn_loss is rebound; anchors, proposal sampling, NMS and the smooth-L1 box term are byte-for-byte identical between arms, so the comparison isolates the loss.
  • mAP is computed from scratch (all-point interpolation, greedy score-ordered matching, one ground-truth box per detection) rather than via pycocotools, so every reported number traces to code in object_detection/common.py.
  • The reported operating point is chosen by argmax F1 on a threshold sweep, while mAP is computed over the full score range with a low inference threshold (0.05). Mixing those two up is a common way to report misleading precision.

5. Version control

This folder is a git repository with the work split across five commits:

git log --oneline

Set your own identity before committing anything further β€” the repo currently carries a placeholder:

git config user.name  "Your Name"
git config user.email "you@students.example.edu"

What is tracked, and why:

Path Tracked Reason
classification_segmentation/, object_detection/, *.py, *.md yes the work
splits/*.csv yes the frozen splits are what make the report reproducible
Results/metrics, Results/plots, Results/overlays yes the evidence behind every number and figure
report.pdf yes the deliverable
Dataset/** no ~850 MB, and redistributing Kaggle data in a repo is not permitted
Results/model_weights/** no large binaries; the brief says to share these via Google Drive

To version checkpoints anyway, use Git LFS rather than committing them raw:

git lfs install && git lfs track "*.pt" && git add .gitattributes

To push to a private remote:

git remote add origin git@github.com:<you>/<repo>.git
git push -u origin main

Keep the repo private until after grading. .gitattributes normalises line endings so run_all.sh stays executable after a round trip through Windows.


6. Verified behaviour

run_tests.py pins down the parts where a silent bug would poison every number in the report. All 66 checks pass:

Area What is asserted
Dependencies every required library is importable, with its version printed, plus a CPU/GPU report
Normalisation XRV input spans [-1024, 1024] and matches xrv.datasets.normalize exactly; student input carries ImageNet statistics; both survive every augmentation policy
Regression guard normalising before augmenting is demonstrated to be destructive, so the correct ordering cannot be reintroduced by accident
Segmentation Dice, IoU, sensitivity, specificity and pixel accuracy on hand-computable masks; HD95 and boundary-F1 on a known shift; empty predictions do not crash
Grad-CAM overlap metrics are 1.0 for saliency entirely inside the lung and 0.0 entirely outside; peak-inside detection; ranking of the inside/outside ratio
Detection IoU, all-point AP, COCO mAP, precision/recall/F1 and TP/FP/FN against hand-computed values; duplicate detections count as false positives; IoU exactly 0.50 is inclusive at the 0.50 threshold and fails at 0.75; size bucketing; zero-detection edge case
Loss ablation the focal and weighted-CE arms leave the smooth-L1 box loss byte-identical to torchvision, focal reduces exactly to cross-entropy at gamma=0, and the swap reverts cleanly
Splits the seed reproduces the split exactly, a different seed does not, splits are disjoint and cover every image, train fraction is 0.700

Separately verified by running: the frozen backbone reports 0 parameters with requires_grad and 6,148 trainable head parameters, and Grad-CAM still produces a non-degenerate map through it.


7. Known gotchas

torchxrayvision collides with model.py. The package vendors third-party code containing an absolute from model.utils import get_norm. Because the assignment requires a file called model.py, Python resolves sys.modules['model'] to your file and the import dies with a confusing No module named 'model.utils'; 'model' is not a package. import_torchxrayvision() in classification_segmentation/model.py hides our entry for the duration of the import and restores it afterwards. Keep that helper if you refactor.

The Kvasir bbox file is misspelled in the official archive β€” it ships as kavsir_bboxes.json, while the assignment writes kvasir_bboxes.json. Both spellings are accepted by the loader.

trainable_backbone_layers is ignored without pretrained weights. torchvision warns and falls back to 5. Expected when you pass --no-pretrained.

torchvision downloads ImageNet backbone weights even when weights=None, because weights_backbone defaults to ImageNet. build_detector sets it explicitly, so --no-pretrained really is from scratch.

Model 1 weights download on first use (30 MB) from the torchxrayvision GitHub release. It is cached in `/.torchxrayvision/`.


8. Submission checklist

  • Folder renamed to your roll number, zipped
  • splits/*_seed16.csv present and unchanged since training started
  • Results/metrics, Results/plots, Results/overlays populated
  • Checkpoints uploaded to Drive, link in the report
  • report_notes.md filled in and python make_report.py re-run
  • No yellow TODO boxes left in report.pdf
  • python run_tests.py reports 66 passed, 0 failed
  • git status is clean and git log reflects your own identity
  • You can explain every preprocessing choice, model choice, metric and result
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support