Alternate Author Name(s)

Mohammad T Khasawneh

Author ORCID Identifier

Sadaf Tabatabaee:

https://orcid.org/0009-0000-0248-4123

Mohammad T. Khasawneh:

https://orcid.org/0000-0002-4302-6943

Document Type

Article

Publication Date

Spring 5-24-2026

Keywords

Keywords: vision-language models; transformers; large language models; radiology report generation; multimodal learning; PRISMA; uncertainty; interpretability; external validation; reproducibility.

Abstract

Transformer vision-language models (VLMs) promise end-to-end automation of radiology reporting and related multimodal tasks. However, evidence remains fragmented across datasets, architectures, evaluation practices, and levels of clinical validation, limiting fair comparison and safe translation into practice. Following PRISMA 2020/PRISMA-S, search engines including PubMed, IEEE Xplore, Web of Science, and Google Scholar were systematically searched for peer-reviewed, English-language studies published between 2019 and 2025 that used paired radiology images and free-text reports. Dual reviewers screened records and extracted data using a locked schema covering datasets, modalities, architectures, training objectives, evaluation metrics, and indicators of clinical readiness. Free-text model descriptions were normalized and clustered using TF-IDF + agglomerative methods to derive an empirical architecture taxonomy, and evidence was synthesized descriptively. From 674 records, 109 studies met the inclusion criteria. Activity surged after 2022 and was dominated by chest radiography, especially MIMIC-CXR and IU-CXR. Three objectives prevailed: full report generation, impression summarization, and diagnostic classification. Evaluation relied mainly on ROUGE, METEOR, and BLEU, whereas entity and clinical metrics, including CheXbert and RadGraph, as well as human reader studies, were uncommon. The derived taxonomy identified recurrent architecture families, including encoder-decoder generators, ViT-to-LLM bridge models, and transformer classifiers, with clear task-architecture coupling. Across the corpus, evidence of external validation, risk-aware evaluation, workflow integration, and reproducibility practices was uneven and generally limited. Transformer-based VLMs have substantially advanced radiologic language generation but remain constrained by modality concentration, evaluation misalignment with clinical correctness, and limited demonstration of real-world readiness. Safe translation will require coordinated progress across data diversity, task-aligned evaluation, uncertainty-aware modeling, and workflow-level validation, supported by standardized, clinically grounded evaluation frameworks and multi-site multimodality evidence.

Comments

Published version of record: https://doi.org/10.1080/24725579.2026.2674247

Publisher Attribution

This is an Accepted Manuscript of an article published by Taylor & Francis in IISE Transactions on Healthcare Systems Engineering on 28 May 2026, available at: https://doi.org/10.1080/24725579.2026.2674247.

Creative Commons License

Creative Commons Attribution-NonCommercial 4.0 International License
This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License

Share

COinS