A Unified Platform for Radiology Report Generation and Clinician-Centered AI Evaluation
medRxiv preprint 2025.07.07.25331018 (2025) · DOI: 10.1101/2025.07.07.25331018
TL;DR
A web platform with Report Generation and Report Evaluation modules supports a clinician-centered Turing test of AI radiology reports; attendings identified AI-generated reports more often than residents (49.9% vs 41.1%).
中文简介:一个包含报告生成与报告评估两个模块的网页平台,用于以临床医生为中心评估 AI 放射报告;在其上开展的图灵测试中,主治医师识别 AI 报告的比例高于住院医师(49.9% 对 41.1%)。
Key points
- An integrated web-based platform with two modules: Report Generation and Report Evaluation.
- Eight participants evaluated 48 anonymized X-ray cases, each with two reports from three comparison groups.
- Attendings outperformed residents in identifying AI-generated reports (49.9% vs 41.1%) and took longer to decide.
Abstract
Generative AI models have demonstrated strong potential in radiology report generation, but their clinical adoption depends on physician trust. In this study, we conducted a radiology-focused Turing test to evaluate how well attendings and residents distinguish AI-generated reports from those written by radiologists, and how their confidence and decision time reflect trust. We developed an integrated web-based platform comprising two core modules: Report Generation and Report Evaluation. Using the web-based platform, eight participants evaluated 48 anonymized X-ray cases, each paired with two reports from three comparison groups: radiologist vs. AI model 1, radiologist vs. AI model 2, and AI model 1 vs. AI model 2. Participants selected the AI-generated report, rated their confidence, and indicated report preference. Attendings outperformed residents in identifying AI-generated reports (49.9% vs. 41.1%) and exhibited longer decision times, suggesting more deliberate judgment. Both groups took more time when both reports were AI-generated. Our findings highlight the role of clinical experience in AI acceptance and the need for design strategies that foster trust in clinical applications.
License: CC BY-NC-ND.
Citation
Zhuoqi Ma, Xinye Yang, Zach Atalay, Andrew Yang, Scott Collins, Harrison X. Bai, Michael Bernstein, Grayson Baird, Zhicheng Jiao. A Unified Platform for Radiology Report Generation and Clinician-Centered AI Evaluation. medRxiv preprint 2025.07.07.25331018 (2025). https://doi.org/10.1101/2025.07.07.25331018
@article{ma2025platform,
title = {A Unified Platform for Radiology Report Generation and Clinician-Centered AI Evaluation},
author = {Ma, Zhuoqi and Yang, Xinye and Atalay, Zach and Yang, Andrew and Collins, Scott and Bai, Harrison X. and Bernstein, Michael and Baird, Grayson and Jiao, Zhicheng},
journal = {medRxiv},
year = {2025},
doi = {10.1101/2025.07.07.25331018}
}
FAQ
What does the platform do?
It combines a Report Generation module and a Report Evaluation module so clinicians can review AI-generated radiology reports side by side with radiologist reports.
How does this relate to the Meta-Radiology paper?
The Meta-Radiology article reports the Turing-test pilot study run on this platform; this preprint describes the platform and the same study.
Where can I try it?
The project page of the evaluation platform is at zachatalay89.github.io/Labsite.
Related work by the authors
- The AI Challenge: A Turing Test Pilot Study of Attendings and Residents in Identifying AI-Generated Content. Meta-Radiology 2025. The peer-reviewed report of the study run on this platform.