Confidence-gated cloud-edge cascade triage via variational risk minimization for medical imaging
Smart Health, vol. 41, 100689 (2026) · DOI: 10.1016/j.smhl.2026.100689
TL;DR
Variational Risk Minimization (VRM) distills a multimodal chest X-ray teacher into an image-only edge model by treating LVLM-generated report variants as Monte Carlo samples; a confidence-gated cloud-edge cascade reaches AUC 0.941 at 103 ms average latency with 20.3% cloud escalation.
中文简介:VRM 把大视觉语言模型生成的多份报告当作潜在临床解读的蒙特卡洛样本,将多模态胸片教师模型蒸馏为只看图像的边缘模型;置信度门控的云边级联在平均 103 毫秒延迟、20.3% 上云比例下达到 AUC 0.941。
Key points
- Emergency chest X-ray triage has a modality gap: reports arrive after triage decisions, yet multimodal foundation models need image and text.
- VRM learns from a variationally marginalized teacher distribution instead of a single teacher target, giving uncertainty-aware supervision when the report is missing.
- Under matched encoder families, VRM outperforms direct fine-tuning baselines, improves calibration, and recovers well from hallucinated supervision.
- Marginalized supervision reduces report-selection instability.
- The confidence-gated cascade reaches AUC 0.941 at 103 ms average latency with 20.3% cloud escalation.
Abstract
Emergency chest X-ray (CXR) triage has a structural modality gap: reports arrive after triage decisions, yet multimodal foundation models require image-text inputs. We present Variational Risk Minimization (VRM), a distillation framework that treats LVLM-generated report variants as Monte Carlo samples of latent clinical interpretations. Rather than distilling from a single teacher target, VRM learns from a variationally marginalized teacher distribution, enabling uncertainty-aware supervision under missing-modality constraints. Under matched encoder families, VRM outperforms direct fine-tuning baselines and improves calibration with strong recovery from hallucinated supervision. Marginalized supervision reduces report-selection instability. In our compact edge-student instantiation, a confidence-gated cascade reaches AUC 0.941 at 103ms average latency with 20.3% cloud escalation, yielding an explicit reliability-latency operating point for cloud–edge clinical workflows.
License: CC BY 4.0 (open access).
Citation
Xinye Yang, Zhusi Zhong, Scott Collins, Michael Bernstein, Grayson Baird, Terrence Healey, Michael Atalay, Mahesh Jayaraman, Xuyu Wang, Zhicheng Jiao. Confidence-gated cloud-edge cascade triage via variational risk minimization for medical imaging. Smart Health, vol. 41, 100689 (2026). https://doi.org/10.1016/j.smhl.2026.100689
@article{yang2026vrm,
title = {Confidence-gated cloud-edge cascade triage via variational risk minimization for medical imaging},
author = {Yang, Xinye and Zhong, Zhusi and Collins, Scott and Bernstein, Michael and Baird, Grayson and Healey, Terrence and Atalay, Michael and Jayaraman, Mahesh and Wang, Xuyu and Jiao, Zhicheng},
journal = {Smart Health},
volume = {41},
pages = {100689},
year = {2026},
doi = {10.1016/j.smhl.2026.100689}
}
FAQ
What problem does this paper solve?
Chest X-ray triage decisions are made before the radiology report exists, so models that need both image and report cannot be used at triage time. The paper trains an image-only model that still benefits from report knowledge.
What is Variational Risk Minimization (VRM)?
A distillation framework that treats several LVLM-generated report variants as Monte Carlo samples of possible clinical interpretations, and trains the student against the teacher's marginalized distribution rather than one teacher target.
How does the cloud-edge cascade work?
A compact student model runs at the edge and handles confident cases; uncertain cases are escalated to the cloud. The reported operating point is AUC 0.941 at 103 ms average latency with 20.3% of cases escalated.
Is the code available?
Yes. The VRM implementation is at github.com/Yangxinyee/vrm-edge-triage, and the related Q-Former distillation work is at github.com/Yangxinyee/Q-DISTILL.
Where was it presented?
It was published in Smart Health (2026) and presented as an oral at IEEE/ACM CHASE 2026.
Related work by the authors
- Reliability Stress Tests and Decision-Time Routing for Chest X-ray Vision-Language Models. CHASE 2026 Workshop. When to escalate from a single VLM to a multi-agent pipeline, measured across 36 configurations.
- Optimization of CNN for Diagnosis on Lung Disease by Lung Segmentation and Rib Suppression. ISCID 2022. Earlier chest X-ray classification work.