Reliability Stress Tests and Decision-Time Routing for Chest X-ray Vision-Language Models
IEEE/ACM CHASE 2026 Workshop (oral) · DOI: 10.1109/CHASE69719.2026.00073
TL;DR
A reliability stress test for chest X-ray vision-language models (CheXagent, MedGemma-4B, MedGemma-27B) across 36 prompt and workflow configurations shows that exact-match accuracy can overstate conservative models; a decision-time router escalates to multi-agent inference only when it helps.
中文简介:针对胸片视觉语言模型(CheXagent、MedGemma-4B、MedGemma-27B)的可靠性压力测试,覆盖 36 种提示词与工作流配置;结果显示仅看精确匹配准确率会高估偏保守的模型,并提出只在有益时才升级到多智能体推理的决策时路由。
Key points
- Two balanced datasets: a private report-backed set and a curated MIMIC subset.
- Three medical VLMs (CheXagent, MedGemma-4B, MedGemma-27B) x three prompt styles x two workflows (single-VLM and multi-agent) = 36 configurations.
- Exact-match accuracy alone can overstate models that default to "Normal" predictions.
- Reliability depends on model family and scale; multi-agent reasoning helps some configurations and hurts others.
- Decision-time routing escalates to multi-agent inference only when beneficial, improving the cost-quality trade-off over fixed workflows.
Abstract
Medical vision-language model (VLM) evaluation is sensitive to workflow design, prompting strategy, and benchmark construction, yet most studies treat these factors in isolation. We introduce a reliability stress test for chest X-ray interpretation built on two balanced datasets (a private report-backed set and a curated MIMIC subset). Three medical VLMs—CheXagent, MedGemma-4B, and MedGemma-27B—are evaluated across three prompt styles and two workflows (single-VLM and multi-agent), producing 36 configurations. We show that exact-match accuracy alone can overstate the effectiveness of conservative models that default to "Normal" predictions. Diagnostic reliability also depends heavily on model family and scale: multi-agent reasoning helps some configurations but hurts others. Building on these observations, we propose a decision-time routing framework that selectively escalates to multi-agent inference only when beneficial, improving the cost–quality trade-off over fixed workflows. Our results highlight the need for evaluation protocols that jointly consider prompt sensitivity, failure-mode diversity, and workflow choice before clinical deployment.
Citation
Xinye Yang, Zhusi Zhong, Scott Collins, Grayson Baird, Xuyu Wang, Zhicheng Jiao. Reliability Stress Tests and Decision-Time Routing for Chest X-ray Vision-Language Models. IEEE/ACM CHASE 2026 Workshop (oral). https://doi.org/10.1109/CHASE69719.2026.00073
@inproceedings{yang2026cxr,
title = {Reliability Stress Tests and Decision-Time Routing for Chest X-ray Vision-Language Models},
author = {Yang, Xinye and Zhong, Zhusi and Collins, Scott and Baird, Grayson and Wang, Xuyu and Jiao, Zhicheng},
booktitle = {2026 IEEE/ACM Conference on Connected Health: Applications, Systems and Engineering Technologies (CHASE)},
year = {2026},
doi = {10.1109/CHASE69719.2026.00073}
}
FAQ
Which models were tested?
CheXagent, MedGemma-4B and MedGemma-27B, each under three prompt styles and two workflows (single-VLM and multi-agent), 36 configurations in total.
Why is exact-match accuracy not enough?
Conservative models that default to "Normal" can score well on exact match while missing abnormal findings, so accuracy alone overstates their reliability.
Does multi-agent reasoning always help?
No. It helps some model and prompt configurations and hurts others, depending on model family and scale.
What is decision-time routing?
A framework that decides per case whether to escalate from a single VLM to multi-agent inference, doing so only when it is expected to help, which improves the cost-quality trade-off over any fixed workflow.
Related work by the authors
- Confidence-gated cloud-edge cascade triage via variational risk minimization for medical imaging. Smart Health 2026. Chest X-ray triage with a confidence-gated escalation from edge to cloud.
- Optimization of CNN for Diagnosis on Lung Disease by Lung Segmentation and Rib Suppression. ISCID 2022. Earlier chest X-ray classification work.