YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Assignment 2 β Topics in AI for Healthcare (Spring 26)
Chest X-ray classification, lung segmentation, Grad-CAM analysis (COVID-19 Radiography Database) and polyp object detection (Kvasir-SEG).
Seed = 16 (last two digits of the roll number). The same seed appears in the code, in every split filename, and in the report.
Before submitting: rename this folder to your full roll number (
mv RollNumber_16 2026XXXX16), set your roll number inreport_notes.md, and zip the result.
1. Setup
Nothing to install if your environment already has the libraries below. Confirm with:
python run_tests.py
Its first section prints every library with its version and fails loudly on anything missing; the remaining sections run 66 self-contained correctness checks that need no dataset.
Libraries
Every entry is genuinely imported by the code β nothing here is speculative.
| Package | Used for | Verified version |
|---|---|---|
torch |
all models, training loops, autograd (Grad-CAM) | 2.14.0 |
torchvision |
DenseNet-121, Faster R-CNN / RetinaNet, transforms.v2 |
0.29.0 |
torchxrayvision |
Model 1 β the frozen chest-X-ray-pretrained DenseNet | 1.5.4 |
numpy |
numerics throughout | 2.4.4 |
pandas |
split CSVs, per-image metric tables | 3.0.2 |
scikit-learn |
stratified splits, confusion matrix, P/R/F1, AUROC | 1.8.0 |
scipy |
HD95 / ASSD distance transforms, ANOVA, Wilcoxon | 1.17.1 |
Pillow |
image and mask IO | 12.1.1 |
opencv-python-headless |
affine warps for joint image/mask augmentation, lung-crop resize | 4.13.0.92 |
matplotlib |
every plot and overlay figure | 3.10.8 |
reportlab |
report.pdf assembly |
4.4.10 |
tqdm |
progress bars β optional, the code degrades gracefully without it | 4.70.0 |
kagglehub |
download_data.py |
β₯0.3 |
torchxrayvision additionally pulls in scikit-image, imageio and requests
at import time; these are listed in requirements.txt so they are not a surprise.
If anything is missing: pip install -r requirements.txt, or
pip install -r requirements-lock.txt for the exact pinned set. For a CUDA build,
install torch / torchvision from https://pytorch.org/get-started/locally/ first.
Datasets
python download_data.py
This pulls both datasets through kagglehub and symlinks them into Dataset/:
| Dataset | Kaggle slug | Size |
|---|---|---|
| COVID-19 Radiography Database | tawsifurrahman/covid19-radiography-database |
~800 MB |
| Kvasir-SEG | debeshjha1/kvasirseg |
~46 MB |
You need Kaggle credentials β ~/.kaggle/kaggle.json, or the KAGGLE_USERNAME /
KAGGLE_KEY environment variables, or an interactive login on first use.
Use --copy if your filesystem does not support symlinks.
Nothing needs to be inside Dataset/: every script also falls back to the kagglehub
cache on its own, and --data-root overrides both.
2. Run everything
bash run_all.sh # full run
bash run_all.sh --quick # tiny run, just to verify the pipeline end to end
Already-trained checkpoints are skipped, so the script is safe to re-run.
Or step by step
Part A β COVID-19 Radiography
cd classification_segmentation
# 2.1 data prep, QC audit, frozen splits, EDA plots
python pre_process.py
# 2.2 Task 1 β Model 1: frozen TorchXRayVision DenseNet
python train.py --task classify --model xrv --epochs 25
python evaluate.py --task classify --ckpt ../Results/model_weights/cls_xrv_seed16.pt
# 2.2 Task 1 β Model 2: student DenseNet-121
python train.py --task classify --model student --init-from imagenet --epochs 25
python evaluate.py --task classify --ckpt ../Results/model_weights/cls_student_imagenet_seed16.pt
# 2.2 the seeded ablation, run identically for both models
python train.py --task classify --model xrv --ablation lung_crop --epochs 25
python train.py --task classify --model student --ablation lung_crop --epochs 25
python evaluate.py --task classify --ckpt ../Results/model_weights/cls_xrv_abl-lung_crop_seed16.pt
python evaluate.py --task classify --ckpt ../Results/model_weights/cls_student_imagenet_abl-lung_crop_seed16.pt
python evaluate.py --compare cls_xrv_seed16 cls_xrv_abl-lung_crop_seed16 \
cls_student_imagenet_seed16 cls_student_imagenet_abl-lung_crop_seed16
# 2.3 Task 2 β lung segmentation
python train.py --task segment --epochs 40
python evaluate.py --task segment --ckpt ../Results/model_weights/seg_unet_seed16.pt
# 2.4 Task 3 β Grad-CAM vs lung mask
python gradcam_analysis.py \
--ckpt-xrv ../Results/model_weights/cls_xrv_seed16.pt \
--ckpt-student ../Results/model_weights/cls_student_imagenet_seed16.pt \
--n-images 12
Part B β Kvasir-SEG
cd object_detection
python pre_process.py # 3.1 + 3.2
python train.py --loss ce --epochs 20 # 3.3 ablation arm 1
python evaluate.py --ckpt ../Results/model_weights/det_fasterrcnn_resnet50_fpn_ce_seed16.pt
python train.py --loss focal --epochs 20 # 3.3 ablation arm 2
python evaluate.py --ckpt ../Results/model_weights/det_fasterrcnn_resnet50_fpn_focal-a0.25g2.0_seed16.pt
python evaluate.py --compare det_fasterrcnn_resnet50_fpn_ce_seed16 \
det_fasterrcnn_resnet50_fpn_focal-a0.25g2.0_seed16
Report
python make_report.py # builds report.pdf from Results/ + report_notes.md
Correctness checks (fast, no dataset, run this before submitting)
python run_tests.py
Sections you have not written yet appear as yellow TODO boxes in the PDF, and the script prints the list of what is still missing.
Useful flags
| Flag | Where | Meaning |
|---|---|---|
--ablation {none,lung_crop,border_mask,gray_input,aug_strong,aug_weak} |
A train | the single preprocessing change |
--imbalance {none,class_weight,sampler,focal} |
A train | class-imbalance handling (default class_weight) |
--init-from {imagenet,scratch} |
A train | Model 2 initialisation |
--loss {ce,focal,weighted_ce} |
B train | the ROI-head loss ablation |
--score-thr |
B evaluate | fixed confidence threshold (default: argmax-F1 from the sweep) |
--data-root |
both pre_process | explicit dataset path |
--workers 0 |
any | if your machine struggles with dataloader workers |
Rough runtimes on one modern GPU: Part A classification ~15 min per run, segmentation
~35 min, Grad-CAM ~2 min, Part B ~30 min per arm. On CPU everything still runs but
expect roughly 20β40Γ longer, so use --quick to check the wiring first.
3. Layout
RollNumber_16/
βββ classification_segmentation/
β βββ common.py config, datasets, transforms, metrics, plot helpers
β βββ pre_process.py indexing, QC audit, frozen splits, EDA
β βββ model.py XRVDenseNet (frozen) | StudentDenseNet | UNet | losses
β βββ train.py classification + segmentation training
β βββ evaluate.py test metrics, ROC, error analysis, overlays, comparisons
β βββ gradcam_analysis.py Grad-CAM for both models + lung-mask overlap metrics
βββ object_detection/
β βββ common.py Kvasir dataset, augmentation, AP/mAP implementation
β βββ pre_process.py annotation checks, frozen split, detection EDA
β βββ model.py Faster R-CNN + the focal / weighted-CE loss swap
β βββ train.py detector training
β βββ evaluate.py mAP suite, threshold sweep, failure analysis
βββ Dataset/{covid19,kvasir}/ datasets (or symlinks to the kagglehub cache)
βββ splits/
β βββ covid_classification_seed16.csv
β βββ covid_segmentation_seed16.csv
β βββ kvasir_detection_seed16.csv
βββ Results/{metrics,plots,overlays,model_weights}/
βββ download_data.py run_all.sh run_tests.py
βββ requirements.txt requirements-lock.txt
βββ report_notes.md make_report.py report.pdf
βββ .gitignore .gitattributes
βββ README.md
Checkpoints in Results/model_weights/ are large β upload them to Google Drive and
put the link in the report, as the assignment instructs.
4. Design decisions you should be ready to defend
These are the choices a viva will probe. Each is implemented and commented in the code.
- Splits are grouped by MD5 before stratifying. An exact duplicate of a training image can therefore never land in test. This dataset is an aggregation of several public sources and does contain repeats.
- No patient identifier exists, so patient-level leakage cannot be fully excluded. Say so in the limitations section rather than claiming a clean split.
- No horizontal flip on chest X-rays. Laterality is anatomically fixed (heart on the patient's left, gastric bubble left) and burned-in L/R markers are real signal. Mirroring produces impossible images. Kvasir frames do get h- and v-flips, because an endoscope view has no canonical orientation.
- Masks are resized with NEAREST and binarised at 0.5; images use BILINEAR. Interpolating a label map invents intermediate values that do not exist.
- Images are 299Γ299, masks are 256Γ256 in this dataset version. They share the same field of view, so both are resized to a common 224. The audit counts these mismatches explicitly rather than silently assuming alignment.
- Model 1 is a genuine linear probe: every backbone parameter has
requires_grad=Falseand the backbone is held ineval()mode so its BatchNorm running statistics stay frozen. It reports ~6.1 K trainable parameters out of ~7.0 M. - Class imbalance (Normal β 10 k vs Viral Pneumonia β 1.3 k) is handled with
inverse-frequency weighted cross-entropy by default;
--imbalance samplerand--imbalance focalare available as alternatives. Balanced accuracy and macro-F1 are reported alongside accuracy for exactly this reason. - Grad-CAM needs
x.requires_grad_(True)for Model 1. With a fully frozen backbone and a non-grad input, autograd builds no graph through the conv stack and activation gradients would beNone. - Boundary metrics (HD95, ASSD, boundary-F1 @2 px) are reported next to Dice. Dice is dominated by the interior of a large lung field and can look healthy while the pleural or diaphragm edge is systematically wrong.
- The detection loss ablation changes exactly one thing. Only
roi_heads.fastrcnn_lossis rebound; anchors, proposal sampling, NMS and the smooth-L1 box term are byte-for-byte identical between arms, so the comparison isolates the loss. - mAP is computed from scratch (all-point interpolation, greedy score-ordered
matching, one ground-truth box per detection) rather than via
pycocotools, so every reported number traces to code inobject_detection/common.py. - The reported operating point is chosen by argmax F1 on a threshold sweep, while mAP is computed over the full score range with a low inference threshold (0.05). Mixing those two up is a common way to report misleading precision.
5. Version control
This folder is a git repository with the work split across five commits:
git log --oneline
Set your own identity before committing anything further β the repo currently carries a placeholder:
git config user.name "Your Name"
git config user.email "you@students.example.edu"
What is tracked, and why:
| Path | Tracked | Reason |
|---|---|---|
classification_segmentation/, object_detection/, *.py, *.md |
yes | the work |
splits/*.csv |
yes | the frozen splits are what make the report reproducible |
Results/metrics, Results/plots, Results/overlays |
yes | the evidence behind every number and figure |
report.pdf |
yes | the deliverable |
Dataset/** |
no | ~850 MB, and redistributing Kaggle data in a repo is not permitted |
Results/model_weights/** |
no | large binaries; the brief says to share these via Google Drive |
To version checkpoints anyway, use Git LFS rather than committing them raw:
git lfs install && git lfs track "*.pt" && git add .gitattributes
To push to a private remote:
git remote add origin git@github.com:<you>/<repo>.git
git push -u origin main
Keep the repo private until after grading. .gitattributes normalises line
endings so run_all.sh stays executable after a round trip through Windows.
6. Verified behaviour
run_tests.py pins down the parts where a silent bug would poison every number
in the report. All 66 checks pass:
| Area | What is asserted |
|---|---|
| Dependencies | every required library is importable, with its version printed, plus a CPU/GPU report |
| Normalisation | XRV input spans [-1024, 1024] and matches xrv.datasets.normalize exactly; student input carries ImageNet statistics; both survive every augmentation policy |
| Regression guard | normalising before augmenting is demonstrated to be destructive, so the correct ordering cannot be reintroduced by accident |
| Segmentation | Dice, IoU, sensitivity, specificity and pixel accuracy on hand-computable masks; HD95 and boundary-F1 on a known shift; empty predictions do not crash |
| Grad-CAM | overlap metrics are 1.0 for saliency entirely inside the lung and 0.0 entirely outside; peak-inside detection; ranking of the inside/outside ratio |
| Detection | IoU, all-point AP, COCO mAP, precision/recall/F1 and TP/FP/FN against hand-computed values; duplicate detections count as false positives; IoU exactly 0.50 is inclusive at the 0.50 threshold and fails at 0.75; size bucketing; zero-detection edge case |
| Loss ablation | the focal and weighted-CE arms leave the smooth-L1 box loss byte-identical to torchvision, focal reduces exactly to cross-entropy at gamma=0, and the swap reverts cleanly |
| Splits | the seed reproduces the split exactly, a different seed does not, splits are disjoint and cover every image, train fraction is 0.700 |
Separately verified by running: the frozen backbone reports 0 parameters with
requires_grad and 6,148 trainable head parameters, and Grad-CAM still produces a
non-degenerate map through it.
7. Known gotchas
torchxrayvision collides with model.py. The package vendors third-party code
containing an absolute from model.utils import get_norm. Because the assignment
requires a file called model.py, Python resolves sys.modules['model'] to your
file and the import dies with a confusing No module named 'model.utils'; 'model' is not a package. import_torchxrayvision() in classification_segmentation/model.py
hides our entry for the duration of the import and restores it afterwards. Keep that
helper if you refactor.
The Kvasir bbox file is misspelled in the official archive β it ships as
kavsir_bboxes.json, while the assignment writes kvasir_bboxes.json. Both spellings
are accepted by the loader.
trainable_backbone_layers is ignored without pretrained weights. torchvision
warns and falls back to 5. Expected when you pass --no-pretrained.
torchvision downloads ImageNet backbone weights even when weights=None, because
weights_backbone defaults to ImageNet. build_detector sets it explicitly, so
--no-pretrained really is from scratch.
Model 1 weights download on first use (30 MB) from the torchxrayvision GitHub
release. It is cached in `/.torchxrayvision/`.
8. Submission checklist
- Folder renamed to your roll number, zipped
-
splits/*_seed16.csvpresent and unchanged since training started -
Results/metrics,Results/plots,Results/overlayspopulated - Checkpoints uploaded to Drive, link in the report
-
report_notes.mdfilled in andpython make_report.pyre-run - No yellow TODO boxes left in
report.pdf -
python run_tests.pyreports 66 passed, 0 failed -
git statusis clean andgit logreflects your own identity - You can explain every preprocessing choice, model choice, metric and result