Author ORCID Identifier
Sadaf Tabatabaee:
https://orcid.org/0009-0000-0248-4123
Mohammad T. Khasawneh:
Document Type
Article
Publication Date
Spring 5-24-2026
Keywords
Keywords: vision-language models; transformers; large language models; radiology report generation; multimodal learning; PRISMA; uncertainty; interpretability; external validation; reproducibility.
Abstract
Transformer vision-language models (VLMs) promise end-to-end automation of radiology reporting and related multimodal tasks. However, evidence remains fragmented across datasets, architectures, evaluation practices, and levels of clinical validation, limiting fair comparison and safe translation into practice. Following PRISMA 2020/PRISMA-S, search engines including PubMed, IEEE Xplore, Web of Science, and Google Scholar were systematically searched for peer-reviewed, English-language studies published between 2019 and 2025 that used paired radiology images and free-text reports. Dual reviewers screened records and extracted data using a locked schema covering datasets, modalities, architectures, training objectives, evaluation metrics, and indicators of clinical readiness. Free-text model descriptions were normalized and clustered using TF-IDF + agglomerative methods to derive an empirical architecture taxonomy, and evidence was synthesized descriptively. From 674 records, 109 studies met the inclusion criteria. Activity surged after 2022 and was dominated by chest radiography, especially MIMIC-CXR and IU-CXR. Three objectives prevailed: full report generation, impression summarization, and diagnostic classification. Evaluation relied mainly on ROUGE, METEOR, and BLEU, whereas entity and clinical metrics, including CheXbert and RadGraph, as well as human reader studies, were uncommon. The derived taxonomy identified recurrent architecture families, including encoder-decoder generators, ViT-to-LLM bridge models, and transformer classifiers, with clear task-architecture coupling. Across the corpus, evidence of external validation, risk-aware evaluation, workflow integration, and reproducibility practices was uneven and generally limited. Transformer-based VLMs have substantially advanced radiologic language generation but remain constrained by modality concentration, evaluation misalignment with clinical correctness, and limited demonstration of real-world readiness. Safe translation will require coordinated progress across data diversity, task-aligned evaluation, uncertainty-aware modeling, and workflow-level validation, supported by standardized, clinically grounded evaluation frameworks and multi-site multimodality evidence.
Publisher Attribution
This is an Accepted Manuscript of an article published by Taylor & Francis in IISE Transactions on Healthcare Systems Engineering on 28 May 2026, available at: https://doi.org/10.1080/24725579.2026.2674247.
Recommended Citation
Tabatabaee, S., & Khasawneh, M. T. (2026). Bridging images and language in radiology: A comprehensive PRISMA systematic review of transformer vision-language models and clinical readiness. IISE Transactions on Healthcare Systems Engineering, 1–23. https://doi.org/10.1080/24725579.2026.2674247
Creative Commons License

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License
Included in
Artificial Intelligence and Robotics Commons, Data Science Commons, Industrial Engineering Commons, Other Operations Research, Systems Engineering and Industrial Engineering Commons, Quality Improvement Commons, Systems Engineering Commons
Comments
Published version of record: https://doi.org/10.1080/24725579.2026.2674247