TL;DR. Before a medical vision-language model (VLM) reads chest X-rays in a clinical workflow, test it the way it will be used: with more than one prompt, more than one workflow, and more than one metric. In our CHASE 2026 workshop paper we ran CheXagent, MedGemma-4B and MedGemma-27B through 36 configurations on two balanced 50-study cohorts. Exact-match accuracy made conservative models look good. A multi-agent pipeline lifted one model's macro F1 from 0.138 to 0.335 and dropped another's from 0.457 to 0.118. The practical answer is to decide per case whether the extra reasoning is worth it. The code, per-case MIMIC-CXR results and tables are in cxr-vlm-routing.
Why a single benchmark number is not enough
Most reports on medical VLMs give one number: accuracy or F1 for one prompt, in one workflow, on one dataset. Deployment changes all three. Someone rewrites the prompt to make outputs shorter, an engineer wraps the model in an agent pipeline, and the hospital's case mix differs from the benchmark.
We wanted to know how much each of those choices moves the result, and in which direction. So instead of proposing a new model we held the weights fixed and varied everything around them.
The setup
- Models. CheXagent-8B, MedGemma-4B-IT and MedGemma-27B-IT, all frozen, greedy decoding.
- Data. Two cohorts of 50 chest X-ray studies each, 25 normal and 25 abnormal. One is a private hospital cohort with labels read from the clinical reports. The other is a curated MIMIC-CXR subset whose labels we audited against the matched reports.
- Labels.
Normalplus 10 findings: atelectasis, cardiomegaly, consolidation, edema, enlarged cardiomediastinum, lung lesion, lung opacity, pleural effusion, pneumonia and pneumothorax. - Prompts. Structured (report the key findings, then choose labels, and don't report a finding without clear evidence), concise (only definite findings, then labels), and labels-only (one line of labels, nothing else).
- Workflows. A single-VLM pass, or a four-stage multi-agent pipeline on the same weights: extract findings, map them to candidate labels from a radiology view, re-check them from an internal-medicine view, and reconcile the two.
Three models, three prompts, two workflows and two datasets give 36 configurations. For each one we report exact match, partial match, macro F1, a hallucination false-positive rate, accuracy on normal cases, partial match on abnormal cases, latency and tokens.
Finding 1: exact match rewards saying "Normal"
Half of each cohort is normal. A model that answers Normal to almost everything gets about half the cases exactly right while missing most disease.
MedGemma-4B shows this clearly. On MIMIC-CXR with the labels-only prompt it reached 52% exact match and 70% partial match in 0.20 seconds per study. Its macro F1 was 0.252, because it covered few pathology classes. MedGemma-27B with the same prompt scored 100% on normal cases and caught at least one correct finding on only 20% of abnormal ones.
What to do: report accuracy on normal and abnormal cases separately, next to macro F1. If normal-case accuracy is near 100% and abnormal partial match is low, the model is defaulting to Normal.
Finding 2: multi-agent reasoning helps some models and hurts others
Adding the four-stage pipeline did not have a consistent effect:
| Configuration | Single-VLM macro F1 | Multi-agent macro F1 |
|---|---|---|
| MIMIC · MedGemma-27B · structured | 0.138 | 0.335 |
| Private · CheXagent · labels-only | 0.457 | 0.118 |
| MIMIC · MedGemma-4B · labels-only | 0.252 | 0.228 |
Paired bootstrap confidence intervals exclude zero for both large shifts: +0.174 [0.059, 0.290] for MedGemma-27B on MIMIC, and +0.295 [0.192, 0.387] in favour of the single pass for CheXagent on the private cohort.
The mechanism is visible in the errors. We saw two opposite failure modes:
- Abnormal-to-normal collapse. The single pass calls an abnormal study
Normal. For MedGemma-27B on MIMIC this happened on 30% of abnormal cases; the extra extraction stages brought it down to 10%. - Over-regularization by consensus. When the stages disagree, the final synthesis falls back to
Normal. CheXagent writes short, sparse outputs, so its extraction stage gave the later stages little to work with. Its collapse rate on the private cohort rose from 10% with a single pass to 44% with the pipeline.
The pipeline also changed hallucinations. CheXagent's single-pass false-positive rate on MIMIC was above 0.35, and the pipeline brought it below 0.13. The MedGemma models' single-pass rates were already below 0.06, and the extra filtering sometimes cost them true positives.
What to do: don't assume an agent pipeline improves a model. Measure it per model, per prompt, and per cohort, and look at what kind of error changed, not just the headline number.
Finding 3: decide per case whether to escalate
If the pipeline helps some cases and hurts others, the fixed choice between "always single" and "always multi-agent" leaves quality or latency on the table. We framed it as a routing decision made at inference time: run the fast single pass, then decide whether this case should go through the multi-agent pipeline.
| Configuration | Strategy | Macro F1 | Latency (s) | Escalated |
|---|---|---|---|---|
| MIMIC · MedGemma-27B · structured | always single | 0.138 | 4.86 | 0% |
| always multi-agent | 0.335 | 6.27 | 100% | |
| oracle | 0.338 | 5.08 | 20% | |
| Private · CheXagent · labels-only | always single | 0.457 | 1.00 | 0% |
| always multi-agent | 0.118 | 2.61 | 100% | |
| oracle | 0.488 | 1.17 | 10% |
The oracle escalates exactly the cases where the pipeline would have helped, so it is a hindsight upper bound, not a method. What it shows is the headroom: on MIMIC, escalating one case in five matches the quality of always running the pipeline for about a sixth of its extra latency (0.22 s instead of 1.41 s per study). On the private cohort, escalating one case in ten beats both fixed workflows.
The paper also tests two simple rules and a small logistic-regression router trained on features of the first pass. With only 50 cases per cohort, the learned router has 5 to 10 positive examples per fold, which is too few to separate the classes reliably. The next step is a router trained on more cases, using signals available at inference time, such as the model's own confidence or disagreement between passes.
A checklist before deploying a chest X-ray VLM
- Test at least three prompt styles; a rewording should not change the ranking of models you're choosing between.
- Report macro F1, normal-case accuracy and abnormal-case partial match together, not exact match alone.
- Count abnormal-to-normal collapses separately; they are the errors that delay care.
- Measure any agent or multi-stage pipeline per model before adopting it.
- Look at hallucinated findings on normal studies, not only missed findings on abnormal ones.
- Use a balanced evaluation set, or report results by class.
- Record latency and tokens next to quality, so the cost of extra reasoning is visible.
- Keep per-case outputs, so a change in the aggregate can be traced to the cases that caused it.
Try it
The repository cxr-vlm-routing (Apache-2.0) has the model wrappers, the prompts and the multi-agent pipeline. It also has per-case outputs for all 18 MIMIC-CXR configurations and the aggregate tables for both cohorts. Rebuilding the tables and checking them against the paper needs no GPU:
python scripts/summarize_formal_results.py
python scripts/build_routing_eval.py
python scripts/verify_paper_numbers.py
To rerun the models, you need your own credentialed copy of MIMIC-CXR-JPG. scripts/prepare_mimic_images.py fetches the 50 studies by ID. The private cohort is not released.
FAQ
Which chest X-ray VLM is best? It depends on the prompt and the workflow. On our MIMIC subset, MedGemma-27B with a structured multi-agent pipeline had the highest macro F1 (0.335). On our private cohort, CheXagent with a labels-only single pass was best (0.457). With 50 cases per cohort, treat these as directions to test, not a leaderboard.
Why is exact-match accuracy misleading for chest X-ray models? With half the cases normal, a model that answers "Normal" by default scores well on exact match while missing disease. Macro F1 and abnormal-case sensitivity expose this.
Do multi-agent pipelines improve medical VLMs? Sometimes. In our tests the same four-stage pipeline raised MedGemma-27B's macro F1 on MIMIC from 0.138 to 0.335 and lowered CheXagent's on a private cohort from 0.457 to 0.118.
What is decision-time routing? Running a fast single pass first, then deciding for each case whether to escalate to a slower multi-agent pipeline. A hindsight oracle shows that escalating 10–20% of cases can match or beat always running the pipeline.
Is the code available? Yes, at github.com/Yangxinyee/cxr-vlm-routing, with per-case MIMIC-CXR results and scripts that check the paper's numbers.