FHIR Server Benchmark Goes Public: Daily Numbers Now Online

US hospital IT teams that watch FHIR tooling have a new public number to argue over. On June 29, 2026, Health Samurai published an open-source benchmark that runs Aidbox, HAPI FHIR, Medplum, and the Microsoft FHIR Server against the same hardware, the same Synthea dataset, and the same k6 workload, with the results posted to a live dashboard that reruns daily. The benchmark was authored by Marat Surmashev, VP of Engineering at Health Samurai, and it is worth saying up front that this is a vendor-run benchmark, since Health Samurai builds Aidbox.

The framing matters because public, reproducible FHIR server numbers are still rare in 2026. Most comparison material for US procurement teams comes from vendor decks or one-off blog posts. A repo that anyone can clone and rerun, with a dashboard that updates every day, is a different shape of evidence.

Same Hardware, Same Data, Four Servers

The setup is a single bare-metal box with 64 CPU cores and 500 GB of RAM. Each FHIR server gets 8 vCPU and 24 GB of RAM, with Medplum running as eight one-vCPU replicas because that matches how it ships. The dataset is Synthea-generated: 1,000 synthetic patient records that expand to roughly 2 million FHIR resources. Aidbox, HAPI, and Medplum back onto PostgreSQL 18; the Microsoft FHIR Server uses SQL Server 2022 Developer Edition, because that is its native pairing.

Load is generated with Grafana k6. The reruns happen daily, and the dashboard at the live dashboard shows the snapshot for any given day.

The Headline CRUD Numbers

The CRUD throughput row is the one most US teams will read first. From the 2026-06-29 snapshot:

  • Aidbox: about 5,212 requests per second
  • HAPI FHIR: about 3,058 requests per second
  • Medplum: about 1,420 requests per second
  • Microsoft FHIR Server: about 440 requests per second

That is roughly a 12x spread between the top and bottom of the four. The P99 latency numbers tell a similar story: Aidbox lands around 91 to 110 ms across create, read, update, and delete, while Microsoft lands above one second on create, update, and delete. HAPI sits between the two extremes.

For more on how US health systems are picking between these servers, the top 5 FHIR terminology servers for US health systems in 2026 writeup covers the adjacent vocabulary side of the same procurement question.

Bundle Import and Search

The benchmark also measures bundle import and FHIR search throughput. Bundle import is the metric that matters for any US team doing data migration from a legacy EHR. The reported numbers: Aidbox at 2,678 resources per second, HAPI at 2,214, Medplum at 764, and Microsoft at 448.

Search is where the ordering shuffles slightly. Aidbox leads at 3,404 RPS, Medplum is second at 1,796, HAPI is third at 1,005, and Microsoft is fourth at 261. The benchmark notes that Medplum does not support composite search, and the Microsoft FHIR Server is very slow on quantity and composite queries specifically.

Storage Footprint Goes the Other Way

Storage size after loading the dataset flips the leaderboard. Microsoft is the smallest at 4.24 GB, Aidbox is next at 6.83 GB, Medplum lands at 11.8 GB, and HAPI is largest at 22.6 GB. The repo notes that HAPI, Medplum, and Microsoft pre-build search indexes on write, while Aidbox ships without default indexes. That makes Aidbox's import faster and its on-disk footprint smaller, but it also means index strategy becomes an operator decision rather than a server default.

What to Watch From Here

The Medplum CTO has already publicly forked the repository, which reads as an optimization response and is exactly the kind of competitive dynamic an open benchmark is supposed to create. Health Samurai has flagged that the current dataset fits comfortably in 500 GB of RAM, and the next post in the series is supposed to test at scale, which is the harder workload to defend against.

For US clinicians and hospital IT leads, the practical read is that a public, reruns-daily FHIR server benchmark is now a normal artefact to ask vendors about. That changes the shape of the procurement conversation more than any single 2026-06-29 number does. For more context, see our digital health newsroom and the commercial vs open-source terminology servers for US payers writeup, which covers an adjacent open-vs-closed evaluation pattern.