
FHIR server performance in production has known patterns. Understanding baseline expectations and common breakage modes prevents surprises.
Baseline throughput (moderate hardware, mid-2026)
| Server | Sustained POST/s | 99p latency | Memory |
|---|---|---|---|
| HAPI JPA 7.x | 1500-1800 | 45 ms | 8 GB |
| Aidbox 2409 | 2500-3000 | 28 ms | 6 GB |
| Medplum 3.x | 3000-3200 | 22 ms | 4 GB |
| Microsoft FHIR (Cosmos) | 1000-1500 | 60 ms | Managed |
| Firely | 2000-2500 | 35 ms | 6 GB |
Read throughput
Generally 3-5x write throughput depending on query complexity. Read caching helps significantly.
Where performance breaks
1. Postgres connection pool saturation. Fixed-size pool at capacity blocks writes. Tune with pgbouncer per HAPI's guide.
2. Terminology $validate-code synchronous. Every write validates a code; slow terminology server cascades. Cache aggressively.
3. Bulk export competing with REST. `$export` consumes DB read capacity; concurrent REST slows.
4. Chained search parameters unindexed. Patient?general-practitioner.name=Smith slow without proper indexes.
5. _include bundle size. _revinclude returning hundreds of resources per Patient consumes memory.
Load testing
1. Use realistic workload mix (not all reads or all writes). 2. Test at 2x expected peak. 3. Test terminology server together, not independently. 4. Test bulk export concurrent with REST. 5. Monitor Postgres metrics during test.
Common performance mistakes
1. Testing single-endpoint performance without workload mix. 2. Missing Postgres index tuning. 3. Cold-start latency ignored. 4. Terminology server undersized. 5. No connection pool tuning.
Performance monitoring
1. Prometheus metrics per resource type. 2. Postgres connection pool metrics. 3. Terminology latency histograms. 4. Bulk export job metrics. 5. Alert on 99p latency spikes.
FHIR server performance is predictable with proper tuning. Sites investing in load testing and Postgres tuning ship reliably; sites without hit performance walls at production scale.
