Meta-Radiology 2025

The AI Challenge: A Turing Test Pilot Study of Attendings and Residents in Identifying AI-Generated Content

Zhuoqi Ma, Xinye Yang, Zach Atalay, Zhusi Zhong, Scott Collins, Harrison X. Bai, Michael A. Bernstein, Ankur Pandey, Eric Dietsche, Joey Z. Gu, Jothika Challapalli, Michael K. Atalay, Mohammad Abubaker-Sharif, Terrance T. Healey, Venkata Paruchuri, Grayson L. Baird, Zhicheng Jiao

Meta-Radiology, 100199 (2025) · DOI: 10.1016/j.metrad.2025.100199

TL;DR

In a radiology Turing test, 4 attendings and 4 residents judged 48 X-ray cases with paired reports; attendings identified AI-generated reports more often than residents (49.9% vs 41.1%) and took longer to decide.

中文简介:一项放射科图灵测试:4 名主治医师和 4 名住院医师评估 48 例胸片的成对报告,主治医师识别 AI 生成报告的比例高于住院医师(49.9% 对 41.1%),且决策时间更长。

Key points

Abstract

Generative Artificial Intelligence (AI) models have demonstrated strong potential in radiology report generation, but their clinical adoption depends on physician trust. In this pilot study, we conducted a radiology-focused Turing test to evaluate how well attendings and residents distinguish AI-generated reports from those written by radiologists, and how their confidence and decision time reflect trust. We developed an integrated web-based platform for report evaluation. Using the web-based platform, eight participants (4 attendings and 4 residents) evaluated 48 anonymized X-ray cases, each paired with two reports from three comparison groups: radiologist vs. AI model 1, radiologist vs. AI model 2, and AI model 1 vs. AI model 2. Participants were asked to select the AI-generated report, rate their confidence, and indicate report preference. Results show that attendings outperformed residents in identifying AI-generated reports (49.9 % vs. 41.1 %) and exhibited longer decision times, suggesting more deliberate judgment. Both groups took more time when both reports were AI-generated. Our findings highlight the role of clinical experience in AI acceptance and the need for design strategies that foster trust in clinical applications.

License: CC BY (open access).

Citation

Zhuoqi Ma, Xinye Yang, Zach Atalay, Zhusi Zhong, Scott Collins, Harrison X. Bai, Michael A. Bernstein, Ankur Pandey, Eric Dietsche, Joey Z. Gu, Jothika Challapalli, Michael K. Atalay, Mohammad Abubaker-Sharif, Terrance T. Healey, Venkata Paruchuri, Grayson L. Baird, Zhicheng Jiao. The AI Challenge: A Turing Test Pilot Study of Attendings and Residents in Identifying AI-Generated Content. Meta-Radiology, 100199 (2025). https://doi.org/10.1016/j.metrad.2025.100199

@article{ma2025radiology,
  title   = {The AI Challenge: A Turing Test Pilot Study of Attendings and Residents in Identifying AI-Generated Content},
  author  = {Ma, Zhuoqi and Yang, Xinye and Atalay, Zach and Zhong, Zhusi and Collins, Scott and Bai, Harrison X. and Bernstein, Michael A. and Pandey, Ankur and Dietsche, Eric and Gu, Joey Z. and Challapalli, Jothika and Atalay, Michael K. and Abubaker-Sharif, Mohammad and Healey, Terrance T. and Paruchuri, Venkata and Baird, Grayson L. and Jiao, Zhicheng},
  journal = {Meta-Radiology},
  pages   = {100199},
  year    = {2025},
  doi     = {10.1016/j.metrad.2025.100199}
}

FAQ

Can radiologists tell AI-generated reports from human ones?

Only partly. In this pilot, attendings identified the AI-generated report 49.9% of the time and residents 41.1% of the time.

How was the study run?

Eight participants used a web-based platform to review 48 anonymized X-ray cases, each paired with two reports, and chose which was AI-generated, rated confidence and stated a preference.

Does experience matter?

Attendings were more accurate and slower than residents, which the authors read as more deliberate judgment.

Related work by the authors