CHASE 2026 WorkshopOral

Reliability Stress Tests and Decision-Time Routing for Chest X-ray Vision-Language Models

Xinye Yang, Zhusi Zhong, Scott Collins, Grayson Baird, Xuyu Wang, Zhicheng Jiao

IEEE/ACM CHASE 2026 Workshop (oral) · DOI: 10.1109/CHASE69719.2026.00073

TL;DR

A reliability stress test for chest X-ray vision-language models (CheXagent, MedGemma-4B, MedGemma-27B) across 36 prompt and workflow configurations shows that exact-match accuracy can overstate conservative models; a decision-time router escalates to multi-agent inference only when it helps.

中文简介:针对胸片视觉语言模型(CheXagent、MedGemma-4B、MedGemma-27B)的可靠性压力测试,覆盖 36 种提示词与工作流配置;结果显示仅看精确匹配准确率会高估偏保守的模型,并提出只在有益时才升级到多智能体推理的决策时路由。

Key points

Abstract

Medical vision-language model (VLM) evaluation is sensitive to workflow design, prompting strategy, and benchmark construction, yet most studies treat these factors in isolation. We introduce a reliability stress test for chest X-ray interpretation built on two balanced datasets (a private report-backed set and a curated MIMIC subset). Three medical VLMs—CheXagent, MedGemma-4B, and MedGemma-27B—are evaluated across three prompt styles and two workflows (single-VLM and multi-agent), producing 36 configurations. We show that exact-match accuracy alone can overstate the effectiveness of conservative models that default to "Normal" predictions. Diagnostic reliability also depends heavily on model family and scale: multi-agent reasoning helps some configurations but hurts others. Building on these observations, we propose a decision-time routing framework that selectively escalates to multi-agent inference only when beneficial, improving the cost–quality trade-off over fixed workflows. Our results highlight the need for evaluation protocols that jointly consider prompt sensitivity, failure-mode diversity, and workflow choice before clinical deployment.

Citation

Xinye Yang, Zhusi Zhong, Scott Collins, Grayson Baird, Xuyu Wang, Zhicheng Jiao. Reliability Stress Tests and Decision-Time Routing for Chest X-ray Vision-Language Models. IEEE/ACM CHASE 2026 Workshop (oral). https://doi.org/10.1109/CHASE69719.2026.00073

@inproceedings{yang2026cxr,
  title   = {Reliability Stress Tests and Decision-Time Routing for Chest X-ray Vision-Language Models},
  author  = {Yang, Xinye and Zhong, Zhusi and Collins, Scott and Baird, Grayson and Wang, Xuyu and Jiao, Zhicheng},
  booktitle = {2026 IEEE/ACM Conference on Connected Health: Applications, Systems and Engineering Technologies (CHASE)},
  year    = {2026},
  doi     = {10.1109/CHASE69719.2026.00073}
}

FAQ

Which models were tested?

CheXagent, MedGemma-4B and MedGemma-27B, each under three prompt styles and two workflows (single-VLM and multi-agent), 36 configurations in total.

Why is exact-match accuracy not enough?

Conservative models that default to "Normal" can score well on exact match while missing abnormal findings, so accuracy alone overstates their reliability.

Does multi-agent reasoning always help?

No. It helps some model and prompt configurations and hurts others, depending on model family and scale.

What is decision-time routing?

A framework that decides per case whether to escalate from a single VLM to multi-agent inference, doing so only when it is expected to help, which improves the cost-quality trade-off over any fixed workflow.

Related work by the authors