Building a payment orchestration layer: routing, retries, failover
Founder & Lead Engineer, RAITHub
A payment orchestration layer sits between your checkout and several payment processors. It picks a processor for each payment by country, currency, cost and success rate, retries or fails over without charging twice, keeps saved cards usable across processors, and turns every provider's webhooks into one event model. Platforms such as Spreedly list 150+ gateway connections, and its vault alone starts at $750 a month.
If you would rather have it built for you, see how RAITHub would build this at the end of this guide.
What is a payment orchestration layer, and do I need one?
It is a piece of your backend, or a platform you rent, that owns one decision: which processor should handle this payment, and what happens if that processor fails. Your checkout calls the layer; the layer calls Stripe, Adyen, a local gateway or a wallet.
You probably need one if at least one of these is true:
- You sell in markets where one processor cannot take the local methods, such as bKash and Nagad in Bangladesh, or wallets elsewhere.
- Your processing fees are large enough that routing a share of volume to a lower-cost processor pays for the engineering.
- An outage at your one processor would stop all revenue, and that risk is no longer acceptable.
- You are already integrating a second processor, and the checkout code is starting to fill with if-statements per provider.
If you sell in one market through one processor, you do not need orchestration yet. A clean payment module with idempotent writes is enough, and it is the foundation an orchestration layer would sit on anyway. That foundation is covered in our payments engineering guide.
How do payment routing rules work?
As two passes: hard rules filter out processors that cannot take the payment, then a score ranks the ones that can. Keep the two separate, so a cost tweak can never route a payment to a processor that does not support it.
| Routing input | Type of rule | Example |
|---|---|---|
| Country of customer or card | Hard filter | Bangladeshi wallets only through the local gateway |
| Currency | Hard filter | Only processors that settle BDT for BDT payments, to avoid conversion |
| Payment method | Hard filter | Wallet payments cannot go to a card-only processor |
| Cost | Score | Percentage fee plus fixed fee, converted to basis points of this amount |
| Recent success rate | Score | Authorisation rate over the last hour or day, per processor and method |
| Health | Hard filter or score | Remove a processor while its error rate is above a threshold (a circuit breaker) |
Fixed fees matter more than people expect. At Stripe's standard US price of 2.9% + 30¢, the 30¢ is 3% of a $10 payment but 0.3% of a $100 one, so the lowest-cost processor can change with the basket size. Score on the real cost of this payment, not on the headline percentage.
How do you retry or fail over without charging the customer twice?
By separating outcomes you know from outcomes you do not, and only ever failing over on the first kind. This is the rule most home-built routers get wrong.
| Outcome from processor | Did money move? | What the router does |
|---|---|---|
| Succeeded | Yes | Stop. Record the provider reference. |
| Declined by the card issuer | No | Stop and show the customer. Do not spray the same card across processors. |
| Not attempted (connection refused, processor marked unhealthy, validation error before execution) | No | Safe to try the next processor. |
| Unknown (timeout, dropped connection, 5xx) | Maybe | Do not fail over. Retry the same processor with the same idempotency key, or resolve by status query or webhook. |
Stripe's own guidance is explicit about the unknown case. Its error-handling docs say that after a network error the client should retry with the same idempotency key and the same parameters until it gets a result, and that a 500 should be treated as indeterminate, because the charge may already have gone out to the card network. If you send that payment to a second processor instead, the customer can be charged by both.
Idempotency keys are the mechanism. Stripe's idempotent requests reference says the server saves the status code and body of the first request for a key and returns the same result for repeats, keys can be up to 255 characters, keys can be pruned after 24 hours, and reusing a key with different parameters returns an error. The same docs suggest deriving keys from something like a cart ID. In an orchestration layer that means one key per order per processor: the same key on every retry to that processor, a different key for a different processor.
Card networks also have rules about retrying declined transactions, and they differ by network and decline reason. Confirm your retry policy with your acquirer; this is general information, not compliance advice.
What does a payment router look like in TypeScript?
A minimal version: filter, score, try processors in order, record every attempt before calling out, and stop on any outcome that is not a definite "not attempted".
type Currency = 'USD' | 'EUR' | 'BDT'
interface ChargeRequest {
orderId: string
amountMinor: number // integer minor units, never floats
currency: Currency
country: string // ISO 3166-1 alpha-2
method: 'card' | 'wallet'
}
type ChargeResult =
| { status: 'succeeded'; providerRef: string }
| { status: 'declined'; code: string } // definitive: issuer said no
| { status: 'not_attempted'; reason: string } // definitive: nothing was processed
| { status: 'unknown' } // timeout or 5xx: may have charged
interface Processor {
id: string
supports(req: ChargeRequest): boolean
charge(req: ChargeRequest, idempotencyKey: string): Promise<ChargeResult>
}
interface Stats { successRate: number; feeBps: number; fixedMinor: number; healthy: boolean }
interface AttemptStore {
findSucceeded(orderId: string): Promise<ChargeResult | null>
record(orderId: string, processorId: string, key: string): Promise<void>
saveResult(key: string, result: ChargeResult): Promise<void>
}
// Higher is better. The weights are a business decision; write them down.
function score(req: ChargeRequest, s: Stats): number {
const costBps = s.feeBps + (s.fixedMinor * 10_000) / req.amountMinor
return s.successRate * 10_000 - costBps
}
export async function route(
req: ChargeRequest,
processors: Processor[],
stats: Map<string, Stats>,
store: AttemptStore,
): Promise<ChargeResult | { status: 'needs_review' } | { status: 'no_processor' }> {
// Our own idempotency: a repeat call for a paid order returns the first result.
const done = await store.findSucceeded(req.orderId)
if (done) return done
const ranked = processors
.filter((p) => p.supports(req) && stats.get(p.id)?.healthy === true)
.sort((a, b) => score(req, stats.get(b.id)!) - score(req, stats.get(a.id)!))
for (const p of ranked) {
const key = req.orderId + ':' + p.id // same key on every retry to p
await store.record(req.orderId, p.id, key)
const result = await p.charge(req, key)
await store.saveResult(key, result)
if (result.status === 'not_attempted') continue // only safe failover
if (result.status === 'unknown') return { status: 'needs_review' }
return result // succeeded or declined
}
return { status: 'no_processor' }
}
Three things this sketch leaves to you. Two checkout requests for the same order must not run the router at the same time, so take a row lock or an advisory lock on the order first. The "needs_review" path needs a background job that retries the same processor with the same key, or queries the status, until the outcome is known. And each processor adapter must classify its own errors correctly into the four outcomes; that mapping is where most of the testing goes.
Why do saved cards stop working when you switch processors?
Because a saved card is usually a token issued by one processor, and a token from one processor generally cannot be charged at another. If all your customers' cards live as Stripe tokens, failing over a returning customer's payment to a second processor is not possible without a vault you control.
There are three ways to keep cards portable:
- An independent vault stores the card once and forwards it to whichever processor you route to. Spreedly's pricing page lists an Independent Vault starting at $750 a month. Primer also describes a central vault for payment credentials.
- Network tokens, issued by the card networks rather than a processor, which both Spreedly and Primer list as features.
- Storing card numbers yourself. This puts your systems in full scope of PCI DSS, the card industry's security standard. For almost every product this is the wrong choice; use a vault.
RAITHub does not build card vaults and makes no PCI DSS claim for its own work. With hosted fields or a provider's vault, raw card numbers never reach the systems we build. Which assessment applies to you is a question for your acquirer or a Qualified Security Assessor; this is general information, not compliance advice.
How do you normalise webhooks from several payment providers?
Map every provider's events into one internal event type at the edge, and let the rest of your system only ever see that type.
interface PaymentEvent {
provider: string // 'stripe', 'sslcommerz', 'bkash', ...
providerEventId: string // unique per provider: dedupe on (provider, providerEventId)
orderId: string
providerRef: string // the provider's payment or transaction ID
type: 'authorized' | 'captured' | 'failed' | 'refunded' | 'disputed'
amountMinor: number
currency: string
receivedAt: Date
raw: unknown // keep the original payload for audits and replays
}
Each provider adapter verifies the signature (or, where a provider has no signatures, calls its validation API), maps the payload into this shape, and stores it under a unique key before acknowledging. Stripe's webhook docs are a good baseline for what every provider can do to you: events can arrive more than once, are not guaranteed to arrive in order, and failed deliveries are retried for up to three days in live mode. Design your order state machine so that a late "authorized" after "captured" changes nothing.
The tests that prove this works, including duplicate and out-of-order delivery, are in testing payments and webhooks.
Should I buy a payment orchestration platform or build a thin layer?
Buy when you need many processors, a portable vault, or routing intelligence across large volume. Build a thin layer when you need two to four processors, including local ones a platform may not cover, and want the routing rules in your own code.
| Option | What the vendor publishes | Choose this when | Trade-off |
|---|---|---|---|
| One processor, no orchestration | For example, Stripe's published standard pricing | One market, one processor covers your methods, an outage is survivable | No failover; you pay one price list |
| Orchestration platform with no-code workflows | Primer describes routing, fallbacks to a secondary processor, a central vault and workflows built "without touching code". It publishes no prices; you book a demo. | Payments or ops teams want to change routing without a deploy | Platform fees; you depend on its connector list |
| Orchestration platform with vault | Spreedly lists 150+ gateway connections, 60+ payment methods, smart retries and PSP-agnostic routing. Its Independent Vault starts at $750 a month; higher tiers are priced by sales. | You need card portability across many processors | Monthly minimum; another vendor in the payment path |
| Custom thin layer | Your own router, adapters and event model | Two to four processors, local wallets or gateways, rules you want in code and under test | You own the adapters, the error mapping and every provider change |
A common middle path is a thin layer for routing and webhooks, with a rented vault for saved cards. You keep control of the logic without taking on card storage.
How long does it take to build a thin orchestration layer yourself?
Our estimate, for an experienced backend developer who knows both providers' APIs: three to six weeks for two processors with filter-and-score routing, idempotent attempts, a normalised webhook model and a status-check job. Each further processor adds an adapter, its error mapping, its webhook mapping and its tests.
The main risk is failing over on an unknown outcome and charging the customer twice. The second is error mapping: one adapter that labels a timeout as "not attempted" turns the router into a double-charge machine. Both are testable before go-live, and both should be.
Why RAITHub for this
- Nine gateways behind one abstraction. PadhAI, the AI tutoring platform built by RAITHub, was designed for 9 payment gateways under one payments abstraction, alongside a 70/20/10 LLM router; routing logic, in both cases, kept in code and under test.
- Local rails in production. TheSkinProof, the founder's own marketplace, takes bKash, Nagad, SSLCommerz and cash on delivery, with 217 API endpoints and 750+ automated tests.
- Stripe inside a SaaS product. PropDesk collects rent through Stripe, under 1,024 automated tests.
- QA-first. Every adapter ships with tests for timeouts, duplicate webhooks and out-of-order events, because those are the failures that move money.
RAITHub has not shipped a regulated or licensed fintech product, does not build card vaults, and makes no compliance claims. We engineer the layer around the processors and vaults you choose.
When you don't need us
- You use one processor in one market. Stay there until a second one is a business need.
- You need dozens of processors or global card portability now. An orchestration platform will get you there faster than any custom build.
- Your product is itself a payment service, such as a payment facilitator or money transmitter. Hire a team that has taken one through licensing and assessment.
How RAITHub would build this
- Scope: a payments module with filter-and-score routing per country, currency, method, cost and success rate, with weights in config.
- Scope: an adapter per processor that classifies every response into succeeded, declined, not attempted or unknown, with idempotency keys per order per processor.
- Scope: one normalised payment event model, with signature verification, deduplication, raw-payload storage and a status-check job for unknown outcomes.
- Scope: health tracking and a circuit breaker per processor, plus a dashboard of routing decisions.
- Scope: integration with your chosen vault or hosted fields, so card numbers stay out of your systems.
Timeline: 6–12 weeks as a backend build, depending on the number of processors and how well each documents its errors.
What you receive: automated tests and CI, including timeout, duplicate and out-of-order scenarios per adapter; handover docs and runbooks for adding a processor and for a processor outage; and full IP under NDA. We sign NDAs and DPAs and work inside your controls; production and financial data stay in your own cloud account, and development uses synthetic data.
Next step: a free 15-minute technical audit, then a written fixed quote. Book the audit, and bring your processor list, markets and current monthly volume by method. More is on the Backend and API development page, the FinTech page and the pricing page.
Frequently asked questions
What is payment orchestration?
Software that sits between your checkout and several payment processors, routes each payment to one of them, handles retries and failover, and presents one API and one event model to the rest of your system.
Is it safe to retry a failed payment on another processor?
Only if the first processor definitely did not process it. On a timeout or server error the outcome is unknown, so retry the same processor with the same idempotency key, or check the status, before trying anywhere else.
How do idempotency keys prevent double charges?
The processor stores the result of the first request for a key and returns it for any repeat. Stripe's docs say keys can be up to 255 characters and may be pruned after 24 hours, so derive the key from the order and processor.
Can I move saved cards from one processor to another?
Not with processor-specific tokens alone. You need an independent vault, network tokens, or a migration arranged between processors. Spreedly lists an Independent Vault starting at $750 a month.
How much does a payment orchestration platform cost?
Primer publishes no prices and asks you to book a demo. Spreedly's vault starts at $750 a month, with higher tiers priced by sales. Processor fees are charged on top in both cases.
How long does it take to build a payment orchestration layer?
By our estimate, three to six weeks for one experienced developer to build a thin layer over two processors. A production build with health tracking, vault integration and full tests is a 6–12 week backend project.
Related posts
Ready to discuss your project?
Book a free 15-minute technical audit with our engineering team.