Blog
Notes on EHR data infrastructure, patient world models, and reliable medical imaging AI.
-
Stress-Testing Chest X-ray Vision-Language Models Before Deployment
· 7 min read
Before a medical vision-language model (VLM) reads chest X-rays in a clinical workflow, test it the way it will be used: with more than one prompt, more than one workflow, and more than one metric. In our CHASE 2026 workshop paper we ran CheXagent, MedGemma-4B and MedGemma-27B through 36 configurations on two balanced 50-study cohorts. Exact-match accuracy made conservative models look good. A multi-agent pipeline lifted one model's macro F1 from 0.138 to 0.335 and dropped another's from 0.457 to 0.118. The practical answer is to decide per case whether the extra reasoning is worth it. The code, per-case MIMIC-CXR results and tables are in cxr-vlm-routing.
-
How to Keep Data Lineage When Building EHR Training Datasets
· 9 min read
An EHR training dataset is only as trustworthy as your ability to answer one question for any row in it: which source rows produced this, and what did the model know at that moment? In practice that means giving every raw row an ID at ingest, deriving every output format from a single canonical event layer, storing event time and information-availability time separately, refusing to guess, making builds reproducible by hash, and validating the artifacts you ship against something independent of them. This post walks through each practice, with examples from building EHR2Trace, an open-source converter from hospital EHR exports to OMOP CDM 5.4 and MEDS.