How FHIR Data Warehouses Transform Digital Health Care in the USA

FHIR Data Warehouses: Design Patterns for Analytics at Scale

FHIR Data Warehouses: Design Patterns for Analytics at Scale

FHIR data warehouses store bulk-exported FHIR resources for analytics. Design patterns shape both performance and operational cost.

Ingestion pipeline

1. Nightly `$export` from FHIR server → NDJSON files in S3. 2. Spark or dbt job flattens NDJSON into warehouse schema. 3. Terminology snapshots joined at ingest time. 4. Feature tables built from warehouse schema.

Warehouse schema patterns

1. Resource-per-table. One table per FHIR resource type. Straightforward but requires cross-table joins for related data.

2. Denormalized wide tables. Patient + Observation + Encounter joined into wide tables. Analytics-friendly but larger storage.

3. Star schema with FHIR fact tables. Traditional data warehouse pattern applied. Encounters or Observations as facts; Patient, Practitioner as dimensions.

4. Data lakehouse (Delta, Iceberg). NDJSON preserved; SQL views project analytics-friendly shapes. Best of both worlds.

Storage sizing (12M patients, mixed resources)

Layer Storage
Raw NDJSON (gzipped) 40-60 GB
Warehouse tables 100-200 GB
Feature stores 10-50 GB
Snapshots + history 500 GB - 1 TB

Common warehouse mistakes

1. Ignoring deleted[] in manifests. MPI-corrupted data. 2. Terminology join at query time. Should join at ingest. 3. Wide tables without indexes. Query scans blow up. 4. No time-based partitioning. Historical query costs balloon. 5. Skipping data quality metrics. Analytics on bad data compounds.

Vendor tooling (mid-2026)

Component Options
FHIR export HAPI, Aidbox, Medplum
Ingest Spark, dbt, Databricks, Fabric
Warehouse BigQuery, Snowflake, Redshift, Fabric
Feature store Feast, dbt marts, custom
BI Tableau, Looker, Power BI

Data quality prerequisites

1. $validate pass rate >97%. 2. Reference integrity >99%. 3. Terminology binding compliance >98%. 4. Complete deleted[] handling.

Cost profile

1. Storage costs scale linearly with retention. 2. Compute costs scale with query complexity. 3. Terminology snapshot storage is small but critical. 4. Feature store maintenance is ongoing engineering.

FHIR data warehouses are a solved discipline. Get the ingestion pipeline, schema pattern, and data quality gates right and analytics scales for years.