Founder & Lead Engineer, RAITHub
RAITHub ships and tests production software. See QA as a Service or talk to us.
When an EHR or FHIR integration breaks, it is usually one of five causes: an expired token or OAuth change, a schema change at the EHR, a rate limit or outage, a silent data-mapping error, or an environment mismatch after a deploy. Read the response first: 401 is auth, 429 is rate limiting, 400 is your payload, 200 with wrong data is a mapping bug. Recover before you re-sync.
This is general engineering information, not compliance advice; keep production and patient data in your own covered cloud and confirm obligations with your adviser. If you would rather have it diagnosed and fixed for you, see how RAITHub would fix this below.
How do I tell what actually broke?
Read the failing response, not the stack trace. The HTTP status and the FHIR error body point straight at the category. FHIR returns errors as a structured OperationOutcome resource, part of the HL7 FHIR specification, so log the whole body, not just the status.
| Status / symptom | Likely cause | First action |
|---|---|---|
| 401 / 403 unauthorized | Expired or revoked token, OAuth scope change | Refresh the token; recheck the registered scopes |
| 429 too many requests | Rate limit hit at peak | Back off and retry; batch and throttle |
| 400 / 422 with OperationOutcome | Payload or profile validation failed | Read the OperationOutcome; compare to the EHR's profile |
| 5xx or timeouts | EHR outage or your request too large | Check the EHR status page; retry idempotently |
| 200 but wrong or missing fields | Silent data-mapping error | Diff a known record end to end |
How do I recover without losing or duplicating data?
The dangerous moment is the re-sync after a failure. If writes are not idempotent, a retry creates a second copy of the same encounter or patient. Make every write safe to retry with an idempotency key, and reconcile before you replay a backlog.
// A retry must not create a second record for the same source event.
async function upsertEncounter(ehrEncounter: FhirEncounter) {
const key = `ehr:${ehrEncounter.id}` // stable key from the source system
return prisma.encounter.upsert({
where: { sourceKey: key },
update: { status: ehrEncounter.status, updatedAt: new Date() },
create: { sourceKey: key, status: ehrEncounter.status /* ... */ },
})
}
Steps to recover safely:
- Stop the retries that are failing, so you are not hammering the EHR or piling up duplicates.
- Find the last good sync point from your own log of processed source IDs.
- Replay from there through an idempotent upsert keyed on the source record's ID.
- Reconcile: count records on both sides and diff a sample, so you prove the catch-up was complete and did not double anything.
Idempotency is the core idea; the general pattern is in FHIR API integration for startups.
How do I stop it breaking again?
- Refresh tokens before they expire and alert on refresh failures, so an expiry never becomes an outage.
- Retry with exponential backoff and a dead-letter queue, so a transient 429 or 5xx recovers itself and a true failure is visible, not lost.
- Pin and test against the EHR's sandbox on a schedule, so a profile or schema change is caught before production.
- Move the sync off the request path into a background job you can observe, retry and replay.
- Keep environment config straight, since sandbox versus production endpoints and client IDs are a common post-deploy break.
The release-safety side of this is in every release breaks something, and the broader recovery mindset is in the HealthTech rescue guide.
Buy, build or hire this?
| Option | Choose this when | Trade-off |
|---|---|---|
| An EHR integration aggregator | You connect to many EHRs and want one API over them | A per-record or subscription cost; you depend on their uptime and coverage |
| Direct integration you maintain | You connect to one or two EHRs and need full control | You own token, schema and rate-limit handling yourself |
| Fix it yourself from the logs | A developer can read a FHIR OperationOutcome and make writes idempotent | Lowest cost; needs someone who knows the EHR's profile and OAuth flow |
| A backend engineer or code rescue | You want it diagnosed, recovered and hardened with tests | An outside dependency; scope the reconciliation and retry tests in |
Doing it yourself is realistic: plan 1 to 3 days if you know the EHR's OAuth flow and your data mapping. The main risk of going alone is replaying a backlog without idempotency and creating duplicate patients or encounters.
Why RAITHub for this
- Idempotent, retry-safe write paths are everyday work. TheSkinProof, the founder's own venture built and run by RAITHub, moves money and stock inside row-locked transactions across 217 API endpoints with 750+ tests.
- Contract-tested APIs and background jobs. PropDesk runs 5 daily automation jobs behind 1,024 tests, the same shape as a resilient EHR sync.
- Honest limits. RAITHub has built and contract-tested REST APIs and a healthcare scheduling app for a client; a FHIR integration is new work, scoped against your EHR partner's version and profiles. It has not shipped a regulated health product and holds no SOC 2 or ISO 27001 certification. Patient data stays in your covered cloud; development uses synthetic or sandbox data.
When you don't need us
- The break is one expired token your own developer can refresh and the writes are already idempotent.
- You use an aggregator and the failure is on their side.
- You have a backend engineer who owns the integration and its tests.
How RAITHub would fix this
- Scope: a free 15-minute call on the symptoms, the EHR and the sync design; you confirm you are authorised to have it diagnosed.
- Diagnose: read the failing responses, classify the cause, and check idempotency and the last good sync point before any replay.
- Recover: replay from the last good point through idempotent upserts against sandbox or synthetic data, then reconcile both sides.
- Harden: token refresh with alerts, backoff and a dead-letter queue, scheduled sandbox tests, and the sync moved to an observable background job.
- What you receive: contract and integration tests in your repository, CI gates, a runbook and full IP; an NDA is standard. Backend work is typically 6 to 12 weeks for a build, less for a scoped recovery.
The work is fixed-price, quoted in writing after the call. See API and backend development, the HealthTech industry page, and the sibling guide on rescuing an inherited HealthTech codebase. To book it, talk to RAITHub about your integration.
General engineering information only; confirm compliance obligations with your adviser. Documentation checked on 11 October 2026.
Frequently asked questions
How do I diagnose a broken EHR or FHIR integration?
Read the failing response. A 401 or 403 is authentication, a 429 is a rate limit, a 400 or 422 with a FHIR OperationOutcome is a payload or profile problem, a 5xx is an outage, and a 200 with wrong data is a silent mapping bug. Log the whole OperationOutcome body, not just the status, so the cause is unambiguous.
How do I re-sync without creating duplicate records?
Make every write idempotent, keyed on the source system's record ID, so a retry updates the existing record instead of creating a second one. Replay from your last known good sync point, then reconcile by counting records on both sides and diffing a sample. Never replay a backlog through non-idempotent writes.
Why did the integration break right after a deploy?
Most often an environment mismatch: sandbox versus production endpoints, a stale client ID or secret, or a config value that did not carry over. Check the environment variables the integration reads, and confirm the token and endpoints match the environment you deployed to.
Should I retry failed EHR requests automatically?
Yes, for transient failures like 429 and 5xx, using exponential backoff and a cap, with a dead-letter queue for requests that keep failing. Do not blindly retry a 400 or 401; those need a fix, not a retry. And only retry writes that are idempotent.
Can RAITHub build a FHIR integration?
RAITHub has built and contract-tested REST APIs and a healthcare scheduling app for a client, which is adjacent experience. A FHIR integration would be new work on your project, scoped against your EHR partner's version and profiles, and tested against their sandbox before production.
Is it safe to test the integration with real patient data?
Test against the EHR's sandbox and synthetic data, not production patient data, and keep production data in your own covered cloud. Treat the integration like any other write path: idempotent, observable and covered by tests. Confirm your data-handling obligations with your adviser.
Related posts
Ready to discuss your project?
Book a free 15-minute technical audit with our engineering team.