Back to BlogIndustry Guides

Subscription Billing and Dunning Done Right

Rupak Amin

Founder & Lead Engineer, RAITHub

10 min read

RAITHub ships and tests production software. See QA as a Service or talk to us.

Subscription billing leaks revenue in two places: webhooks processed twice or missed, and dunning that gives up too early or never stops. Treat every billing webhook as applied-exactly-once; model failed payments as a dunning state machine (retry, warn, grace, downgrade) with timed steps; and test both. Get these right and recurring revenue and account access stay correct even when payments fail, which they always do.

This is general engineering guidance, not financial or compliance advice; confirm tax and dunning-notice rules with your adviser. If you would rather have it built and tested for you, see how RAITHub would build this below.

Where does subscription billing actually go wrong?

Not in charging a card that works, but in everything around it: the webhook that tells you the charge happened, and the recovery when it does not. A webhook arrives twice and you grant two months; a webhook is missed and a paying customer loses access; a card fails and dunning either stops too soon (you lose a recoverable customer) or never stops (you keep emailing someone who churned). Billing is an event-driven system, and it fails at the events, not the charge.

FailureWhat happensThe fix
Duplicate webhookAn event is processed twice; access or credit is double-grantedApplied-exactly-once handling keyed to the event ID
Missed webhookA real charge or cancellation is never appliedReconcile against the provider, do not trust webhooks alone
Dunning too shortA recoverable failed payment is abandonedA retry schedule timed to how cards actually recover
Dunning never endsA churned customer is emailed forever, or keeps accessA terminal downgrade state the machine must reach

How do you handle billing webhooks safely?

Treat every webhook as applied-exactly-once: they arrive duplicated, retried and out of order. Record each provider event by its ID before acting on it, so a repeat is a no-op, and make the action idempotent so even a race cannot double-apply it. The full treatment, including testing against replayed and out-of-order events, is in testing payments and webhooks end to end, and the underlying principle is in idempotency in API design. A common, costly bug is a webhook handler that returns the wrong status so the provider keeps retrying a success, covered in the troubleshooting post Stripe webhook not firing or returning 405.

-- Record every provider event by its id; process each one once.
CREATE TABLE billing_events (
  provider_event_id text PRIMARY KEY,     -- the provider's event id, unique
  type        text NOT NULL,
  processed_at timestamptz,
  created_at  timestamptz NOT NULL DEFAULT now()
);

-- On receipt: insert the id (ON CONFLICT DO NOTHING). If it was already there,
-- the event is a duplicate -> acknowledge and do nothing.

What is dunning, and why is it a state machine?

Dunning is how you recover a failed subscription payment: a sequence of retries and messages, ending either in recovery or a deliberate downgrade. It is a state machine because each failed invoice moves through defined stages with timed waits, and it must always reach a terminal state, never loop forever. A workable shape: payment failed → retry scheduled → warned → grace period → past due → downgraded or recovered. Each transition is driven by a scheduled job, not by anyone watching, so recovery does not depend on a person.

CREATE TABLE dunning (
  id            uuid PRIMARY KEY DEFAULT gen_random_uuid(),
  subscription_id bigint NOT NULL REFERENCES subscriptions(id),
  state         text NOT NULL DEFAULT 'payment_failed',
                -- payment_failed | retrying | warned | grace | past_due | recovered | downgraded
  attempt       integer NOT NULL DEFAULT 0,
  next_action_at timestamptz NOT NULL,        -- when the scheduled job acts next
  created_at    timestamptz NOT NULL DEFAULT now()
);

CREATE INDEX dunning_due ON dunning (next_action_at) WHERE state NOT IN ('recovered','downgraded');

How should retries be timed?

To match how failed cards actually recover: many fail for a temporary reason (insufficient funds, a limit) and succeed on a later attempt, so retries spread over days recover more than retries crammed into hours. A common pattern is a handful of attempts over a week to two weeks, with a warning email before access is affected and a clear final step. The provider often offers smart retry logic; use it where it fits, but keep your own state machine as the source of truth for what the customer can access, because entitlement is your decision, not the provider's. Each attempt updates attempt and next_action_at; the scheduled job acts on whatever is due. Billing and failed-payment notice requirements vary by jurisdiction; this is general information, so confirm them with your adviser.

How do you keep access correct during dunning?

Entitlement (what the customer can do) must follow the subscription's real state, and must be restored the instant a retry succeeds. Keep a single source of truth for entitlement that reads the subscription and dunning state, so a recovered payment immediately returns access and a reached downgrade immediately limits it. Do not scatter access checks that each guess at billing status, that is how a paying customer stays locked out after recovery, or a churned one keeps access. Churn you can prevent with engineering (failed-payment recovery, access that reflects reality) is covered in reducing SaaS churn with engineering, and billing models themselves in usage-based billing for SaaS.

How do you test billing and dunning?

Test the two failure classes directly: duplicate events and the dunning path.

import { describe, it, expect } from 'vitest'
import { handleWebhook, runDunning, stateOf } from './billing'

describe('billing + dunning', () => {
  it('processes a duplicate webhook only once', async () => {
    const event = paymentSucceeded({ id: 'evt_1', sub: 5 })
    await handleWebhook(event)
    await handleWebhook(event)                 // same id, retried
    expect(await monthsGranted(5)).toBe(1)     // not doubled
  })

  it('reaches a terminal state after retries are exhausted', async () => {
    const d = await seedDunning({ sub: 5, state: 'payment_failed' })
    for (let i = 0; i < 6; i++) { await failRetry(d.id); await runDunning() }
    expect(['downgraded', 'recovered']).toContain(await stateOf(d.id))
  })
})

Add a test that a successful retry restores access immediately, one that reconciles a missed webhook against the provider, and one that out-of-order events (a success after a failure) leave the right final state.

Buy, build or hire?

OptionChoose this whenThe catch
A billing provider's built-in dunning (Stripe Billing and similar)Standard plans and the provider's retry and email flow fit youYou still own entitlement; relying on the provider for access decisions drifts from your real state
A billing platform (a subscription-management vendor)Complex plans, tax and invoicing you do not want to buildYou fit their model and fees, and still wire entitlement into your app correctly
Custom dunning on top of a providerYour entitlement rules are specific, or you need dunning you control and can testYou own the state machine and the tests that keep revenue and access correct

How long does it take to build yourself, and what is the risk?

For an experienced developer building idempotent webhooks, a dunning state machine and entitlement on top of a provider, our estimate is 2 to 4 weeks, or 4 to 6 weeks with reconciliation, proration and multiple plans. The main risk is treating webhooks as reliable: they are not, so a system that trusts them without idempotency and reconciliation will both double-grant and silently miss charges. Test with replayed and out-of-order events, not just the happy path. Notice and tax rules vary by jurisdiction; this is general information, so confirm them with your adviser.

Why RAITHub for this

  • Stripe billing in production. RAITHub built PropDesk with Stripe rent collection and idempotent webhooks, across 1,024 tests, and builds Stripe subscriptions, metering, invoices and dunning as standard SaaS work.
  • Idempotent money events. Applied-exactly-once handling is everyday work, which is what keeps billing from double-granting under retries.
  • Entitlement done once. Access reads the real subscription state, so recovery and downgrade take effect immediately.

When you don't need us

  • The provider's dunning fits. Standard plans and its retry flow may be enough, as long as entitlement is wired correctly.
  • You are pre-revenue. Do not over-build billing before you charge.
  • You only need the model. The event table and dunning machine above are a fair start for your own developer.

How RAITHub would build this

  • Idempotent webhooks: every provider event recorded by ID and applied exactly once, with reconciliation against the provider so nothing is missed.
  • Dunning state machine: timed retries matched to how cards recover, a warning, a grace period and a terminal downgrade, driven by a scheduled job.
  • Entitlement: a single source of truth that restores or limits access the instant the subscription state changes.
  • Tests: duplicate-event, out-of-order, terminal-state and access-restore suites gated in CI.

Timeline: adding this to an existing SaaS is typically 2 to 4 weeks; a full billing backend with proration and reconciliation fits the 4 to 6 week fixed-scope range. See SaaS development, QA as a Service and the FinTech industry page. A related build is an expense-management SaaS.

You receive: the billing and dunning code, tests gated in CI, handover runbooks, and full IP under NDA. Tax and notice rules are confirmed with your adviser, not claimed by RAITHub.

Frequently asked questions

Why does subscription billing lose revenue even when charges work?

Because it fails around the charge: a webhook processed twice double-grants, a missed webhook loses a paying customer access, and dunning either abandons a recoverable payment or never stops. Billing is event-driven, so it fails at the events, not the successful charge.

How do I handle billing webhooks safely?

Treat every webhook as applied-exactly-once. Record each provider event by its ID before acting, so a repeat is a no-op, make the action idempotent, and reconcile against the provider so a missed event is still caught. Webhooks arrive duplicated, retried and out of order, so never trust them alone.

What is dunning, and why model it as a state machine?

Dunning is how you recover a failed subscription payment through timed retries and messages, ending in recovery or a deliberate downgrade. It is a state machine because each failed invoice moves through defined stages and must always reach a terminal state, never loop forever.

How should failed-payment retries be timed?

To match how cards recover: many fail temporarily and succeed on a later attempt, so a handful of retries spread over one to two weeks recovers more than retries crammed into hours. Warn the customer before access is affected, and keep a clear final step.

How do I keep account access correct during dunning?

Keep a single source of truth for entitlement that reads the subscription and dunning state, so a successful retry restores access immediately and a reached downgrade limits it at once. Scattered access checks that each guess at billing status are how customers get wrongly locked out or keep access after churning.

Should I use the provider's dunning or build my own?

Use the provider's retries and emails where they fit, but keep your own state machine as the source of truth for what the customer can access, because entitlement is your decision. Build custom dunning when your access rules are specific or you need behaviour you can test and control.

subscription billingdunningfailed paymentsstripe webhooksfintechsaas

Ready to discuss your project?

Book a free 15-minute technical audit with our engineering team.