Back to BlogArchitecture & Engineering

Sending Webhooks to Your SaaS Customers: Retries, Signing and Logs

Rupak Amin

Founder & Lead Engineer, RAITHub

12 min read

To send webhooks to your SaaS customers reliably, write each event to an outbox table in the same transaction as the change, deliver from a background queue, sign every request with HMAC-SHA256 using the Standard Webhooks headers, retry failures with backoff for about a day, refuse internal network addresses, and show customers a delivery log with a replay button.

If you would rather have webhook delivery built for you, see how RAITHub would build this below.

What are outgoing webhooks, and why do SaaS customers ask for them?

An outgoing webhook is an HTTP POST your system sends to a URL the customer registers, each time something they care about happens: an invoice is paid, a ticket is closed, a user is invited. It replaces polling, where the customer calls your API every minute to ask "anything new?"

Customers ask for webhooks because they want to connect your product to their CRM, accounting tool or internal workflows in near real time. For you, webhooks cut API load from polling and make integrations easier to build. They pair with a public API: the webhook says "invoice inv_123 changed", and the API lets the customer fetch the full, current invoice. The API side is covered in designing a public API for your SaaS.

What does a reliable webhook delivery system look like?

Five parts, and every one of them is needed. Sending an HTTP request inside your web request handler, the first version most teams write, fails on all five.

PartWhat it doesWhat goes wrong without it
OutboxRecords the event in the same database transaction as the changeThe change commits but the event is lost if the process crashes before sending
Queue and workersDelivers in the background, with a timeout per attemptA slow customer endpoint slows down your own users' requests
SigningLets the customer prove the request came from you and was not alteredAnyone who finds the URL can send fake events
Retries and disablingRetries failures with backoff; pauses endpoints that keep failingA short customer outage means lost events, or endless retries to dead URLs
Logs and replayShows each attempt, status and response; lets customers resendEvery missed event becomes a support ticket only you can investigate

How should webhooks be signed?

Follow the Standard Webhooks specification, an open spec written so that senders and receivers can share libraries. It defines three headers: webhook-id, a unique message ID; webhook-timestamp, Unix seconds; and webhook-signature. The signed content is msg_id.timestamp.payload, signed with HMAC-SHA256. The secret is base64-encoded and prefixed with whsec_, and the signature is sent as v1, followed by the base64 signature. Several space-separated signatures are allowed, which is how you rotate a secret without downtime.

import { createHmac, timingSafeEqual } from 'node:crypto'

// Sender side: sign exactly the bytes you will send.
export function sign(secret: string, msgId: string, timestamp: number, body: string): string {
  const key = Buffer.from(secret.replace(/^whsec_/, ''), 'base64')
  const signed = msgId + '.' + timestamp + '.' + body
  return 'v1,' + createHmac('sha256', key).update(signed).digest('base64')
}

export function webhookHeaders(secrets: string[], msgId: string, body: string) {
  const timestamp = Math.floor(Date.now() / 1000)
  return {
    'content-type': 'application/json',
    'webhook-id': msgId,
    'webhook-timestamp': String(timestamp),
    // During rotation, sign with old and new secret; the receiver accepts either.
    'webhook-signature': secrets.map((s) => sign(s, msgId, timestamp, body)).join(' '),
  }
}

// Receiver side (put this in your docs): constant-time compare, reject stale timestamps.
export function verify(secret: string, headers: Record<string, string>, body: string, toleranceSec = 300): boolean {
  const ts = Number(headers['webhook-timestamp'])
  if (!Number.isFinite(ts) || Math.abs(Date.now() / 1000 - ts) > toleranceSec) return false
  const expected = Buffer.from(sign(secret, headers['webhook-id'], ts, body).slice(3))
  return (headers['webhook-signature'] ?? '').split(' ').some((entry) => {
    const [version, sig] = entry.split(',')
    const given = Buffer.from(sig ?? '')
    return version === 'v1' && given.length === expected.length && timingSafeEqual(given, expected)
  })
}

The spec asks receivers to use a constant-time comparison and to check that the timestamp is within a tolerance, to stop replay attacks; five minutes is a common choice. Sign the raw body bytes, and tell customers to verify before parsing JSON: re-serialising can change whitespace or key order and break the signature. That one mistake is behind many "signature mismatch" tickets, the same problem receivers hit in Stripe webhooks that do not fire.

How often should failed webhooks be retried?

Retry with growing gaps for about a day, then stop and mark the message failed. The Standard Webhooks spec gives an example schedule: immediately, then after 5 seconds, 5 minutes, 30 minutes, 2 hours, 5 hours, 10 hours, 14 hours, 20 hours and 24 hours. It counts any 2xx status as success and suggests a request timeout of 15 to 30 seconds.

  • Retry on timeouts, connection errors, 5xx and 429. Respect a Retry-After header if the customer sends one.
  • Do not retry most 4xx responses forever; a 401 or 404 usually means a configuration problem that more attempts will not fix. Retry a few times, then fail and notify.
  • Add jitter to each delay so that one customer's outage does not produce a synchronised wave of retries when it recovers.
  • Disable dead endpoints. After every message to an endpoint has failed for several days, pause it and email the customer's admin. Resume and replay from the dashboard.
  • Do not promise ordering. Retries make strict order impossible to guarantee. Include a timestamp or version in the payload and tell customers to fetch the latest state from the API when order matters.

Because a message can arrive more than once, the spec tells receivers to use webhook-id as an idempotency key. Keep the same ID on every retry of the same message.

How do you store and queue outgoing webhooks?

Use a transactional outbox: insert the event in the same transaction as the business change, and let workers pick it up. If the transaction rolls back, no event exists; if it commits, the event cannot be lost.

CREATE TABLE webhook_messages (
  id            text PRIMARY KEY,              -- 'msg_' + random; becomes webhook-id
  tenant_id     uuid NOT NULL,
  event_type    text NOT NULL,                 -- 'invoice.paid'
  payload       jsonb NOT NULL,
  created_at    timestamptz NOT NULL DEFAULT now()
);

CREATE TABLE webhook_attempts (
  message_id    text NOT NULL REFERENCES webhook_messages(id),
  endpoint_id   uuid NOT NULL,
  attempt       int  NOT NULL,
  next_try_at   timestamptz NOT NULL,
  status        text NOT NULL DEFAULT 'pending', -- pending, succeeded, failed
  response_code int,
  response_body text,                            -- truncated, for the customer's log
  duration_ms   int,
  PRIMARY KEY (message_id, endpoint_id, attempt)
);

-- Workers claim due work without blocking each other.
SELECT message_id, endpoint_id, attempt
FROM webhook_attempts
WHERE status = 'pending' AND next_try_at <= now()
ORDER BY next_try_at
LIMIT 50
FOR UPDATE SKIP LOCKED;

FOR UPDATE SKIP LOCKED lets several workers poll the same table without picking the same row. At higher volumes you can move delivery to a dedicated queue, but keep the outbox as the source of truth. Keep message bodies for a set retention period, such as 30 days, and say so in your docs; payloads can contain personal data.

How do you stop customers using webhooks to attack your network?

Treat every customer-supplied URL as hostile. Server-side request forgery (SSRF), where your server is tricked into calling an internal address, is API7 in the OWASP API Security Top 10 (2023), and a webhook URL field is a direct way in.

  • HTTPS only in production, with certificate validation on.
  • Resolve the hostname yourself and block private ranges: 127.0.0.0/8, 10.0.0.0/8, 172.16.0.0/12, 192.168.0.0/16, 169.254.0.0/16 (cloud metadata lives here), and their IPv6 equivalents. Connect to the IP you checked, so a DNS change between check and request cannot redirect you.
  • Do not follow redirects, or re-check every hop.
  • Send from a separate egress path with no route to your internal services, and publish its IP addresses so customers can allow-list them.
  • Cap the response you read and store, and time out every attempt.

What should the customer-facing webhook dashboard include?

Enough that customers can debug their own endpoint without opening a ticket: add and test endpoints; choose event types per endpoint; reveal and rotate the signing secret; a log of every message with each attempt's status code, response snippet and duration; a replay button per message and "replay all failed since" per endpoint. Record endpoint and secret changes in your audit log, as described in SaaS audit log design.

Do-it-yourself estimate: 1–2 weeks for the outbox, workers, signing, retries, SSRF guard, dashboard and tests, if you already run background jobs. The main risk is the parts that only show up under failure: a customer endpoint that hangs, a queue that backs up, or a signature mismatch you cannot reproduce without logs.

Buy, build or hire?

Hosted webhook services handle delivery, retries, signing and a customer portal for you. As one published example, Svix pricing lists a free plan with 50,000 messages a month, a Basic plan from $20 a month and a Professional plan from $490 a month, with additional messages at $0.0001 each (checked 7 October 2026).

OptionExamplesChoose this whenWatch out for
Hosted webhook serviceWebhooks-as-a-service platforms with an embeddable customer portalYou want delivery, retries and a portal in days, and volumes fit the pricingEvent payloads pass through another processor; add it to your DPA and security answers
Open-source library or templateStandard Webhooks reference libraries for signing, plus your job queueYou have developers and a job runner, and want to control data residencyLibraries cover signing; retries, SSRF, logs and the portal are still yours
Custom buildThe outbox design aboveWebhooks are core to your product, or payloads cannot leave your infrastructureOperating it: monitoring, backlogs and disabled endpoints become your on-call work
Hire a team to build itRAITHub or another studioCustomers are waiting and your team is on the core productGet the runbooks and failure tests in the handover

How do you test outgoing webhooks?

  • Signing: verify your own signatures with the published receiver code, including during secret rotation with two signatures.
  • Failure modes: point an endpoint at a test server that returns 500, hangs past the timeout, returns 410 and closes the connection, and check each retry path.
  • SSRF: try localhost, private ranges, the metadata address, IPv6 loopback and a redirect to each; all must be refused.
  • Outbox integrity: roll back a transaction and check no event exists; kill a worker mid-delivery and check the message is retried, not lost.
  • Load: generate a burst of events for one tenant and check other tenants' deliveries are not delayed.

The receiving side has its own checklist in testing payments and webhooks, and the wider approach is in the API testing guide.

Why RAITHub for this

  • Webhooks from the receiving side, many times over. PropDesk collects rent through Stripe and PadhAI runs 9 payment gateways; TheSkinProof, the founder's own venture rather than a client, integrates bKash, Nagad and SSLCommerz. Each involves handling a provider's payment callbacks or webhooks correctly, which is the mirror image of sending them.
  • Back ends with real test coverage. TheSkinProof has 217 API endpoints and 750+ tests; PropDesk has 1,024 tests.
  • Multi-tenant isolation. BlockEstate is multi-tenant, so per-tenant endpoints and secrets are familiar ground.

RAITHub has not published a case study of an outgoing webhook platform; the guidance above is engineering practice, not a client result.

When you don't need us

  • You send a few event types to a handful of customers. A hosted webhook service's free tier will cover you.
  • You only need email notifications. Webhooks are for machines; do not build them for people.
  • You want developers placed in your team. RAITHub does fixed-scope builds and dedicated teams, not staff augmentation.

How RAITHub would build this

  • Event catalogue: named, versioned event types with documented payloads and examples.
  • Delivery engine: transactional outbox, workers with timeouts, Standard Webhooks signing, retries with jitter, and automatic endpoint disabling.
  • Safety: an SSRF guard, a dedicated egress path with published IPs, and payload retention rules.
  • Customer dashboard: endpoints, event filters, secret rotation, delivery logs and replay.
  • Or integration instead: if a hosted service fits better, wiring it to your outbox so events are still never lost.

Timeline: added to an existing back end, this fits the 6–12 week backend and API range alongside other work; inside a new SaaS build it can be part of the 4–6 week fixed scope.

You receive: failure-mode tests in CI, receiver sample code for your docs, handover docs and runbooks, and full IP assigned to you under NDA.

Next step: a free 15-minute technical audit, then a written fixed quote. See the API and backend development service, or book the audit.

Frequently asked questions

How do I sign webhooks I send to customers?

Compute an HMAC-SHA256 over the message ID, timestamp and raw body, joined by dots, with a per-endpoint secret, and send it in a webhook-signature header. That is the Standard Webhooks format, and customers can verify it with existing libraries.

How long should I retry a failed webhook?

About a day with growing gaps is common; the Standard Webhooks spec's example schedule ends at 24 hours. After that, mark the message failed, notify the customer and let them replay it.

Should webhooks guarantee delivery order?

No. Retries make strict ordering impractical. Include a timestamp or version in each payload and let customers fetch the current state from your API when order matters.

What is a webhook outbox?

A database table where you record each event in the same transaction as the change that caused it. Background workers deliver from it, so an event can never be lost between a commit and a send.

How do I prevent SSRF through webhook URLs?

Allow HTTPS only, resolve the hostname and block private, loopback and link-local addresses, connect to the address you checked, refuse redirects, and send from a network path with no access to internal services.

Should I build webhooks or use a hosted service?

Use a hosted service when you need it working quickly and payloads may pass through a third party. Build it when webhooks are central to the product or data must stay in your infrastructure.

WebhooksOutgoing webhooksStandard WebhooksHMAC signingRetriesSSRFSaaS

Ready to discuss your project?

Book a free 15-minute technical audit with our engineering team.