5. Conclusion and future work
This thesis presented CXR-HIEU, a unified vision–language model that performs chest-X-ray findings generation, impression generation, and visual question answering with a single RAD-DINO + projection + Vicuna-7B (LoRA) backbone, guided by an explicit Positive/Negative/Uncertain abnormality signal and trained with a parameter-efficient two-stage schedule. On the MIMIC-CXR test split the model is competitive with established systems on the standard report-generation metrics (BLEU-1 0.364, BLEU-4 0.103, METEOR 0.161, ROUGE-L 0.292), its abnormality classifier reaches a macro-F1 of 0.328 (≈ 0.42 on the labels it can predict), and it achieves a VQA exact match of 0.306, markedly stronger on closed-ended questions. These results show that a frozen-encoder, frozen-LLM design adapted only through a small projection and LoRA adapters can drive three chest-X-ray tasks at once on modest hardware.
Several directions would strengthen the work:
- Stronger abnormality classifier. The clearest lever: adopt a META-CXR-style design (multi-encoder fusion, class-balancing or focal loss, contrastive/uncertainty objectives) so the predicted PNU string is more reliable and propagates fewer errors into generation and VQA.
- Clinical validation. Assess the generated reports with expert radiologists for factual correctness, beyond the lexical and classifier metrics used here.
- Impression generation. Investigate and close the gap on impression (dedicated decoding budget, task-specific tuning, or a true end-to-end findings→impression cascade).
- Scale and backbones. Train on more data and views (multi-image studies) and evaluate stronger LLMs (e.g. Llama-3).
- Cross-dataset validation. Evaluate on IU X-ray and CheXpert to measure generalisation beyond MIMIC-CXR.
References
E. Çallı, E. Sogancioglu, B. van Ginneken, K. G. van Leeuwen, K. Murphy. Deep Learning for Chest X-ray Analysis: A Survey. Medical Image Analysis, vol. 72, 2021.
D. Edirisinghe, W. Nimalsiri, M. Hennayake, D. Meedeniya, G. Lim. Chest X-Ray Report Generation Using Abnormality Guided Vision Language Model (META-CXR). IEEE Access, vol. 13, 2025.
C. Pellegrini, E. Özsoy, B. Busam, N. Navab, M. Keicher. RaDialog: A Large Vision-Language Model for Radiology Report Generation and Conversational Assistance. arXiv:2311.18681, 2023.
Z. Wang, L. Liu, L. Wang, L. Zhou. R2GenGPT: Radiology Report Generation with Frozen LLMs. Meta-Radiology, 2023.
Y. Li, Z. Wang, Y. Liu, L. Wang, L. Liu, L. Zhou. KARGEN: Knowledge-enhanced Automated Radiology Report Generation using Large Language Models. arXiv:2409.05370, 2024.
S. Bannur et al. MAIRA-2: Grounded Radiology Report Generation. arXiv:2406.04449, 2024.
Z. Wang, L. Liu, L. Wang, L. Zhou. METransformer: Radiology Report Generation by Transformer with Multiple Learnable Expert Tokens. CVPR, 2023.
Z. Huang, X. Zhang, S. Zhang. KiUT: Knowledge-injected U-transformer for Radiology Report Generation. CVPR, 2023.
F. Pérez-García et al. RAD-DINO: Exploring Scalable Medical Image Encoders Beyond Text Supervision. arXiv:2401.10815, 2024.
J. Li, D. Li, S. Savarese, S. Hoi. BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models. ICML, 2023.
W.-L. Chiang et al. Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90% ChatGPT Quality. 2023.
E. J. Hu et al. LoRA: Low-Rank Adaptation of Large Language Models. ICLR, 2022.
T. Dettmers, A. Pagnoni, A. Holtzman, L. Zettlemoyer. QLoRA: Efficient Finetuning of Quantized LLMs. NeurIPS, 2023.
A. Smit et al. CheXbert: Combining Automatic Labelers and Expert Annotations for Accurate Radiology Report Labeling Using BERT. EMNLP, 2020.
J. Irvin et al. CheXpert: A Large Chest Radiograph Dataset with Uncertainty Labels and Expert Comparison. AAAI, 2019.
A. E. W. Johnson et al. MIMIC-CXR-JPG, a Large Publicly Available Database of Labeled Chest Radiographs. arXiv:1901.07042, 2019.
T. Zhang, V. Kishore, F. Wu, K. Q. Weinberger, Y. Artzi. BERTScore: Evaluating Text Generation with BERT. ICLR, 2020.
S. Banerjee, A. Lavie. METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments. ACL Workshop, 2005.
J. J. Lau, S. Gayen, A. Ben Abacha, D. Demner-Fushman. A Dataset of Clinically Generated Visual Questions and Answers about Radiology Images (VQA-RAD). Scientific Data, 2018.
B. Liu, L.-M. Zhan, L. Xu, L. Ma, Y. Yang, X.-M. Wu. SLAKE: A Semantically-Labeled Knowledge-Enhanced Dataset for Medical Visual Question Answering. IEEE ISBI, 2021.
X. Hu et al. Expert Knowledge-Aware Image Difference Graph Representation Learning for Difference-Aware Medical Visual Question Answering (Medical-Diff-VQA / MIMIC-Diff-VQA). ACM SIGKDD (KDD), 2023.
S. Bae et al. EHRXQA: A Multi-Modal Question Answering Dataset for Electronic Health Records with Chest X-ray Images (incl. MIMIC-CXR-VQA). NeurIPS Datasets & Benchmarks, 2023.