How to Keep Data Lineage When Building EHR Training Datasets

Xinye Yang · · 9 min read

TL;DR. An EHR training dataset is only as trustworthy as your ability to answer one question for any row in it: which source rows produced this, and what did the model know at that moment? In practice that means giving every raw row an ID at ingest, deriving every output format from a single canonical event layer, storing event time and information-availability time separately, refusing to guess, making builds reproducible by hash, and validating the artifacts you ship against something independent of them. This post walks through each practice, with examples from building EHR2Trace, an open-source converter from hospital EHR exports to OMOP CDM 5.4 and MEDS.


Why lineage matters more for training data than for analytics

Most EHR pipelines were designed for cohort studies and dashboards, where a small error shifts an estimate. Training data for patient world models, clinical agents and offline reinforcement learning fails differently. A model learns whatever regularity the data contains, including the ones the pipeline introduced. A diagnosis that appears a few hours before it was actually entered, a cohort label that leaks in as a condition, a lab value copied onto the wrong timestamp: each becomes a shortcut the model will happily learn, and none of them makes the pipeline crash.

Lineage is what lets you find these problems after the fact. If every published event links back to the source rows it came from, you can audit a suspicious prediction down to the export that caused it. Without lineage, a clean-looking dataset is a claim you cannot check.

1. Treat the raw export as read-only and give every row an ID at ingest

The first transformation is where lineage is usually lost, so it is where it should start.

If the source is a normalized relational database, flatten it into one table per logical source first, so that lineage stays row-level instead of pointing at a join.

In EHR2Trace, ingest writes one Parquet file per logical source with the original values and a source_row_id, and the preparation scripts write a manifest of input and output hashes.

2. Derive every output format from one canonical event layer

It is tempting to write an OMOP ETL and a separate MEDS ETL. Two pipelines means two sets of decisions about identity, time and units, and they will drift.

A better structure is:

raw export (read-only)
    ↓ ingest                one lineage record per row
source/                     original values + source_row_id
    ↓ identity, canonical   patient identity, time semantics, values, units, dedup
canonical/                  the single event store
    ↓                ↓
omop/               meds/
    ↓ validate      checks on what was written, with lineage

Every decision about what an event is happens once, in the canonical layer. Each canonical event keeps links to its source rows, and both targets inherit them. In EHR2Trace every canonical, OMOP and MEDS record traces to a source_row_id, and OMOP carries an etl_audit lineage table.

3. Store event time and information-availability time separately

This is the single most important practice for training data.

A lab result has a time the specimen was collected and a later time the result became available. A diagnosis code has a clinical time and a time it was entered, sometimes days later at billing. If you store only one timestamp, you will eventually put information in front of the model before a clinician could have seen it. That is temporal leakage, and it inflates offline metrics while producing a model that fails in deployment.

Keep both clocks on every event, and let the downstream episode builder decide which one to respect. Watch for layouts that invite the wrong choice: one export I worked with repeated every ancillary result once per index study and carried the study's date on each repeated row. Reading that column as the result's own time is exactly the mistake the layout suggests.

4. Keep orders, dispensing and administration as distinct records

A medication order is not evidence the patient received the drug. Pharmacy dispensing is closer; administration records are closer still. Collapsing them into one "drug exposure" event is convenient for analytics and wrong for a model that is supposed to learn what happened to the patient. Keep them as separate event types with their own times and lineage, and let the consumer decide how to combine them.

5. Refuse to guess

Many silent errors start as a reasonable default: assume the source timezone, assume a missing birth year from an age, map an unknown unit to the most common one. Each default is a decision the data never made.

Two mechanisms help:

6. Make builds reproducible by hash

If you cannot rebuild the same dataset, you cannot tell whether a change in model behaviour came from the model or the data.

A side benefit is incremental rebuilds. When outputs are content-addressed, a rerun reuses whatever already exists for identical inputs.

7. Validate what ships, against something independent

This is the lesson that cost me the most to learn.

A validation suite that reads each output in isolation will pass on a surprising number of broken datasets. When I injected 28 silent faults drawn from real incidents into a fixture, every one is now detected, but 15 of the 28 had gotten past the checks at the time they happened. Nine of those fifteen were found only by re-reading three finished conversions, about 300 million events that had already been published, validated and reported clean.

The six misses found by the fault-injection experiment itself had a common cause: each was a check looking at the wrong artifact.

Silent fault Why a single-artifact check missed it
Anchor dates used as event times The configuration was checked, but the data already carried anchor dates as clinical times.
Subject IDs the identity map never issued The identity map was internally consistent; nothing joined it to the events.
A declared text column that published nothing Every check read what was published; none compared it to what the configuration said would be published.
A MEDS layer built without the vocabulary its OMOP layer used Each target validated on its own; nothing compared the two.

The checks added in response all have the same shape: check the artifact that ships, against something derived independently of it, such as another artifact, the dataset's own declaration, or a reference table. Cross-layer checks compare merged rows against their source rows, values and units against reference ranges, published deaths against the deaths the canonical layer holds, and MEDS concepts against OMOP concepts.

Two more rules make validation honest:

Also note where the faults lived: 21 of the 28 never reach an OMOP database, so CDM-level data-quality tooling alone cannot see them. They are in the canonical layer, the manifest, the identities or the MEDS shards.

8. Keep language models out of the data path

LLMs are useful for proposing terminology mappings or interpreting column names. They should not parse dates, numbers or table structure, invent concept IDs, or write to any output layer. In EHR2Trace the rule is enforced by a test: a model may propose, and a person confirms. Accepted decisions are compiled into a mapping registry, which is the only thing that turns a source string into a concept ID. Whether model-assisted ranking actually helps reviewers is itself measured against confirmed decisions rather than assumed.

A checklist

Trying it

EHR2Trace implements all of the above. It runs on the openly licensed MIMIC-IV demonstration subset (100 patients) without credentials, and a single command traces one patient through every layer:

.venv/bin/ehr2trace trace -d datasets/mimiciv.yaml --patient <key>

The repository is at github.com/Yangxinyee/ehr2trace (Apache-2.0). The fault catalogue with every injected fault and the incident behind it is in docs/FAULT_CATALOGUE.md. The companion paper, EHR2Trace: Auditable EHR Data Infrastructure for Patient World Models and Clinical Agents, is forthcoming.

FAQ

What is data lineage in an EHR dataset? The ability to trace any published record back to the exact source rows, configuration and code version that produced it.

Is OMOP or MEDS better for training models? They serve different purposes: OMOP is a relational common data model built for observational research, while MEDS is a minimal event-stream format built for machine learning. Deriving both from one canonical layer avoids choosing and keeps them consistent.

What is temporal leakage in EHR data? Using information at a time before it was actually available, for example a diagnosis code entered days after the visit but timestamped at the visit. Storing information-availability time separately prevents it.

Can I use EHR2Trace with my hospital's export? Yes. Everything dataset-specific lives in one YAML configuration; a new export needs a new configuration and no core code changes. The repository contains no patient data.