TL;DR. An EHR training dataset is only as trustworthy as your ability to answer one question for any row in it: which source rows produced this, and what did the model know at that moment? In practice that means giving every raw row an ID at ingest, deriving every output format from a single canonical event layer, storing event time and information-availability time separately, refusing to guess, making builds reproducible by hash, and validating the artifacts you ship against something independent of them. This post walks through each practice, with examples from building EHR2Trace, an open-source converter from hospital EHR exports to OMOP CDM 5.4 and MEDS.
Why lineage matters more for training data than for analytics
Most EHR pipelines were designed for cohort studies and dashboards, where a small error shifts an estimate. Training data for patient world models, clinical agents and offline reinforcement learning fails differently. A model learns whatever regularity the data contains, including the ones the pipeline introduced. A diagnosis that appears a few hours before it was actually entered, a cohort label that leaks in as a condition, a lab value copied onto the wrong timestamp: each becomes a shortcut the model will happily learn, and none of them makes the pipeline crash.
Lineage is what lets you find these problems after the fact. If every published event links back to the source rows it came from, you can audit a suspicious prediction down to the export that caused it. Without lineage, a clean-looking dataset is a claim you cannot check.
1. Treat the raw export as read-only and give every row an ID at ingest
The first transformation is where lineage is usually lost, so it is where it should start.
- Never modify the raw export. Read it; write everything else somewhere else.
- Assign a stable
source_row_idto every row as it is parsed, and keep the original values next to any normalized ones. - Record a manifest: input files with their hashes, rows added or dropped by any preparation step, and every delivered file the pipeline did not read.
If the source is a normalized relational database, flatten it into one table per logical source first, so that lineage stays row-level instead of pointing at a join.
In EHR2Trace, ingest writes one Parquet file per logical source with the original values and a source_row_id, and the preparation scripts write a manifest of input and output hashes.
2. Derive every output format from one canonical event layer
It is tempting to write an OMOP ETL and a separate MEDS ETL. Two pipelines means two sets of decisions about identity, time and units, and they will drift.
A better structure is:
raw export (read-only)
↓ ingest one lineage record per row
source/ original values + source_row_id
↓ identity, canonical patient identity, time semantics, values, units, dedup
canonical/ the single event store
↓ ↓
omop/ meds/
↓ validate checks on what was written, with lineage
Every decision about what an event is happens once, in the canonical layer. Each canonical event keeps links to its source rows, and both targets inherit them. In EHR2Trace every canonical, OMOP and MEDS record traces to a source_row_id, and OMOP carries an etl_audit lineage table.
3. Store event time and information-availability time separately
This is the single most important practice for training data.
A lab result has a time the specimen was collected and a later time the result became available. A diagnosis code has a clinical time and a time it was entered, sometimes days later at billing. If you store only one timestamp, you will eventually put information in front of the model before a clinician could have seen it. That is temporal leakage, and it inflates offline metrics while producing a model that fails in deployment.
Keep both clocks on every event, and let the downstream episode builder decide which one to respect. Watch for layouts that invite the wrong choice: one export I worked with repeated every ancillary result once per index study and carried the study's date on each repeated row. Reading that column as the result's own time is exactly the mistake the layout suggests.
4. Keep orders, dispensing and administration as distinct records
A medication order is not evidence the patient received the drug. Pharmacy dispensing is closer; administration records are closer still. Collapsing them into one "drug exposure" event is convenient for analytics and wrong for a model that is supposed to learn what happened to the patient. Keep them as separate event types with their own times and lineage, and let the consumer decide how to combine them.
5. Refuse to guess
Many silent errors start as a reasonable default: assume the source timezone, assume a missing birth year from an age, map an unknown unit to the most common one. Each default is a decision the data never made.
Two mechanisms help:
- Blockers before conversion. Before reading any data, list the questions only the data owner can answer, such as the source timezone or the reference date an age is measured from, and refuse to run until they are answered in configuration. EHR2Trace's
inspectstep prints these and exits non-zero while any remain open. - Quarantine during conversion. Anything that cannot be resolved with certainty goes to a quarantine table or a human review queue, with its lineage, instead of being coerced into a plausible value.
6. Make builds reproducible by hash
If you cannot rebuild the same dataset, you cannot tell whether a change in model behaviour came from the model or the data.
- Pin the environment, including the version of the data standard (MEDS's schema is a data contract).
- Address outputs by content: input hash, configuration hash and code version. Same input, configuration and mapping version should give the same output hash.
- Test it: run the build at several concurrency settings and assert that they produce one digest.
A side benefit is incremental rebuilds. When outputs are content-addressed, a rerun reuses whatever already exists for identical inputs.
7. Validate what ships, against something independent
This is the lesson that cost me the most to learn.
A validation suite that reads each output in isolation will pass on a surprising number of broken datasets. When I injected 28 silent faults drawn from real incidents into a fixture, every one is now detected, but 15 of the 28 had gotten past the checks at the time they happened. Nine of those fifteen were found only by re-reading three finished conversions, about 300 million events that had already been published, validated and reported clean.
The six misses found by the fault-injection experiment itself had a common cause: each was a check looking at the wrong artifact.
| Silent fault | Why a single-artifact check missed it |
|---|---|
| Anchor dates used as event times | The configuration was checked, but the data already carried anchor dates as clinical times. |
| Subject IDs the identity map never issued | The identity map was internally consistent; nothing joined it to the events. |
| A declared text column that published nothing | Every check read what was published; none compared it to what the configuration said would be published. |
| A MEDS layer built without the vocabulary its OMOP layer used | Each target validated on its own; nothing compared the two. |
The checks added in response all have the same shape: check the artifact that ships, against something derived independently of it, such as another artifact, the dataset's own declaration, or a reference table. Cross-layer checks compare merged rows against their source rows, values and units against reference ranges, published deaths against the deaths the canonical layer holds, and MEDS concepts against OMOP concepts.
Two more rules make validation honest:
- A check with nothing to examine should report skipped, not passed.
- Known gaps belong in configuration as declared thresholds, so they are reported with their measurement instead of being tuned away.
Also note where the faults lived: 21 of the 28 never reach an OMOP database, so CDM-level data-quality tooling alone cannot see them. They are in the canonical layer, the manifest, the identities or the MEDS shards.
8. Keep language models out of the data path
LLMs are useful for proposing terminology mappings or interpreting column names. They should not parse dates, numbers or table structure, invent concept IDs, or write to any output layer. In EHR2Trace the rule is enforced by a test: a model may propose, and a person confirms. Accepted decisions are compiled into a mapping registry, which is the only thing that turns a source string into a concept ID. Whether model-assisted ranking actually helps reviewers is itself measured against confirmed decisions rather than assumed.
A checklist
- Raw export is read-only; all outputs live outside the source and outside the code repository.
- Every raw row has a
source_row_id; a manifest records input hashes and unread files. - One canonical event layer; OMOP, MEDS or any other target is derived from it.
- Event time and information-availability time are stored separately.
- Orders, dispensing and administration are distinct event types.
- Open questions block the run; uncertain values go to quarantine or review.
- Outputs are content-addressed; rebuilds are tested for identical digests.
- Validation checks shipped artifacts against independent references, and reports skips.
- Faults you have actually hit are injected into a fixture and gated in CI.
- Language models propose; people decide; nothing model-written reaches an output layer.
Trying it
EHR2Trace implements all of the above. It runs on the openly licensed MIMIC-IV demonstration subset (100 patients) without credentials, and a single command traces one patient through every layer:
.venv/bin/ehr2trace trace -d datasets/mimiciv.yaml --patient <key>
The repository is at github.com/Yangxinyee/ehr2trace (Apache-2.0). The fault catalogue with every injected fault and the incident behind it is in docs/FAULT_CATALOGUE.md. The companion paper, EHR2Trace: Auditable EHR Data Infrastructure for Patient World Models and Clinical Agents, is forthcoming.
FAQ
What is data lineage in an EHR dataset? The ability to trace any published record back to the exact source rows, configuration and code version that produced it.
Is OMOP or MEDS better for training models? They serve different purposes: OMOP is a relational common data model built for observational research, while MEDS is a minimal event-stream format built for machine learning. Deriving both from one canonical layer avoids choosing and keeps them consistent.
What is temporal leakage in EHR data? Using information at a time before it was actually available, for example a diagnosis code entered days after the visit but timestamped at the visit. Storing information-availability time separately prevents it.
Can I use EHR2Trace with my hospital's export? Yes. Everything dataset-specific lives in one YAML configuration; a new export needs a new configuration and no core code changes. The repository contains no patient data.