Usage-Based Billing for SaaS: Metering, Stripe Meters and Invoices
Founder & Lead Engineer, RAITHub
Usage-based billing works when every billable event is written once to your own database, rolled into hourly totals and sent to the billing system with an idempotency key, so a retry never charges twice. On Stripe, meter events must be timestamped within the past 35 days (Stripe docs), so a sender that falls behind silently loses revenue.
If you would rather have it built for you, see how RAITHub would build this below.
This guide is the engineering layer under the pricing decision. If you are still choosing between flat, per-seat, usage and hybrid pricing, start with SaaS billing models explained; for subscriptions, checkout and webhooks, see how to add Stripe billing to your SaaS. Here we cover what neither of those goes deep on: event design, idempotent ingestion, aggregation windows, late events, credits, invoice previews, plan limits and overage alerts.
What does a usage-based billing system actually need?
Five parts, in this order: a usage event you own, an ingestion path that accepts each event exactly once, an aggregation step, a sender that reports totals to the billing system, and reconciliation that proves the invoice matches your records.
| Part | What it does | What goes wrong without it |
|---|---|---|
| Usage event | One row per billable action: tenant, metric, quantity, when it happened, an idempotency key | You cannot answer "why is my bill higher?" with evidence |
| Idempotent ingestion | A unique constraint on tenant plus key, so a retried request writes nothing new | Client retries and queue redeliveries double-bill |
| Aggregation window | Rolls events into per-tenant, per-metric, per-hour totals | One API call per click hits rate limits |
| Sender | Reports each total once, with a stable identifier, and records the result | A crash between "sent" and "saved" bills twice or not at all |
| Reconciliation | Compares your totals with the invoice line before and after finalization | Under-billing goes unnoticed for months |
The rule that holds it together: your own table is the source of truth, and the billing system is a downstream copy. That is also what lets you enforce plan limits in real time, which matters because Stripe's own docs note that Billing Meters "only reconciles usage at invoice time" (Stripe usage-based billing).
Should you use Stripe Billing meters, Metronome, Orb or Lago?
For a new Stripe integration, Stripe itself now points you at Metronome. Its docs call Metronome, now part of Stripe, "Stripe's primary usage-based billing platform, recommended for all new integrations", while basic usage-based billing on the Billing Meters API "remains fully supported for existing integrations" (Stripe docs). Stripe lists prepaid credits, enterprise commits, dimensional pricing and real-time usage visibility as reasons to choose Metronome, and full compatibility with Connect, Checkout, Adaptive Pricing or Workflows as reasons to stay on Billing Meters.
| Option | Pricing (as published) | Choose it when | Watch out for |
|---|---|---|---|
| Stripe Billing meters | Pay-as-you-go Billing is "0.7% of Billing volume" plus payment fees (Stripe Billing pricing) | You already bill on Stripe, have a few simple metrics, and need Checkout or Connect compatibility | Usage reconciles at invoice time; you build live usage display and limits yourself |
| Metronome (part of Stripe) | Through Stripe; see the comparison in Stripe's docs | New usage integrations, credits and commits, enterprise contracts, high event volume | Stripe notes limited support for some Stripe products such as Connect and Checkout |
| Orb | Custom pricing on Core, Advanced and Enterprise tiers, based on billings, events and a platform fee (Orb pricing) | Pricing changes often and finance wants warehouse sync, Salesforce or NetSuite | A sales conversation before you see a number |
| Lago | Open source under AGPLv3 to self-host, or Lago Cloud and Premium on request (Lago on GitHub, Lago pricing) | You want to run billing on your own infrastructure or inspect the rating engine | You operate it: upgrades, backups and on-call are yours |
Whichever you choose, the pipeline in this guide stays the same: own the events, aggregate, send idempotently, reconcile. Only the sender's last call changes.
Buy, build or hire?
| Route | Examples | Choose this when | The catch |
|---|---|---|---|
| Off-the-shelf billing platform | Metronome, Orb, Lago Cloud | Usage pricing is core to the business, with contracts, commits and many metrics | You still have to emit clean, idempotent events from your product; no vendor does that part |
| No-code or template | A meter and metered price created in the Stripe Dashboard, usage sent from a scheduled script; or self-hosted Lago | One metric, low volume, and you want to test whether customers accept usage pricing | No live usage view, no limits, and scripts that fail quietly |
| Custom pipeline on a billing API | Your own usage tables and sender, with Stripe or Metronome doing rating and invoices | You need real-time limits, customer-facing usage pages and evidence for every invoice line | It is real engineering: tests, monitoring and reconciliation, not a weekend job |
| Fully custom rating and invoicing | Your own pricing engine and invoice generation | Rarely; only when no vendor can express your contract or tax rules | You now own tax, dunning, proration and invoice law. Usually not worth it |
How should you design a billable usage event?
Record the business action, not the HTTP request, and give each one a key that the caller generates and reuses on retry. A good event has:
- tenant_id: who pays. In multi-tenant SaaS this is the organization, not the user.
- metric: a stable name such as api_requests or ai_tokens. Stripe ties one event name to one meter (Stripe meter configuration), so renaming later means a new meter.
- quantity: a number. Stripe accepts decimals, and if a cycle's total is negative it reports the invoice quantity as 0 (Stripe docs).
- occurred_at and received_at: when it happened and when you heard about it. You need both to handle late events.
- idempotency_key: for example the job ID, message ID or request ID. Not a random value made on the server, which defeats the point.
Keep dimensions (region, model, plan) few. Stripe accepts up to 10,000 unique dimension combinations per meter per hour, and up to 100 per customer on a meter; events past those limits are invalid (Stripe docs).
Choose the aggregation formula on purpose. Stripe meters offer three, and the meter cannot be changed after configuration apart from its display name (Stripe docs):
| Formula | Bills on | Typical metric |
|---|---|---|
| Sum | Total of all values in the period | Tokens, GB transferred, minutes |
| Count | Number of events in the period | Documents processed, API calls |
| Last | Most recent value reported | Active seats or stored GB at period end |
How do you make usage ingestion idempotent?
With a database constraint, not application logic. The usage table's primary key is the tenant plus the caller's idempotency key, so the second copy of an event is simply ignored. A second table holds hourly batches, and each batch's ID becomes the identifier sent to the billing system.
CREATE TABLE usage_events (
tenant_id uuid NOT NULL REFERENCES tenants(id),
idempotency_key text NOT NULL,
metric text NOT NULL,
quantity numeric(20,6) NOT NULL CHECK (quantity >= 0),
occurred_at timestamptz NOT NULL,
received_at timestamptz NOT NULL DEFAULT now(),
batch_id uuid,
PRIMARY KEY (tenant_id, idempotency_key)
);
CREATE INDEX usage_events_unbatched ON usage_events (received_at) WHERE batch_id IS NULL;
CREATE INDEX usage_events_period ON usage_events (tenant_id, metric, occurred_at);
CREATE TABLE meter_batches (
id uuid PRIMARY KEY DEFAULT gen_random_uuid(),
tenant_id uuid NOT NULL REFERENCES tenants(id),
metric text NOT NULL,
hour_start timestamptz NOT NULL,
total numeric(20,6) NOT NULL,
status text NOT NULL DEFAULT 'pending'
CHECK (status IN ('pending', 'sent', 'failed', 'too_old')),
attempts int NOT NULL DEFAULT 0,
last_error text,
sent_at timestamptz
);
-- Ingest: a retried event is a no-op.
INSERT INTO usage_events (tenant_id, idempotency_key, metric, quantity, occurred_at)
VALUES ($1, $2, $3, $4, $5)
ON CONFLICT (tenant_id, idempotency_key) DO NOTHING;
Where possible, write the usage row in the same transaction as the billable action. If the action rolls back, so does the charge. In a shared-schema SaaS, put the usage table under the same row-level security as everything else, so a tenant's usage page can never show another tenant's numbers; the pattern is in the Postgres row-level security guide.
How do aggregation windows and late events work?
Roll unbatched events into one batch per tenant, metric and UTC hour, but only events received more than a couple of minutes ago, so in-flight inserts are not split. Hourly batches keep you far below Stripe's live-mode limit of 1,000 meter event calls per second, and its rule of one concurrent call per customer per meter (Stripe docs).
BEGIN;
SET LOCAL TimeZone = 'UTC'; -- Stripe uses UTC hour boundaries
SELECT pg_advisory_xact_lock(4210); -- one batching worker at a time
WITH new_batches AS (
INSERT INTO meter_batches (tenant_id, metric, hour_start, total)
SELECT tenant_id, metric, date_trunc('hour', occurred_at), sum(quantity)
FROM usage_events
WHERE batch_id IS NULL AND received_at < now() - interval '2 minutes'
GROUP BY 1, 2, 3
RETURNING id, tenant_id, metric, hour_start
)
UPDATE usage_events e
SET batch_id = nb.id
FROM new_batches nb
WHERE e.batch_id IS NULL
AND e.received_at < now() - interval '2 minutes'
AND e.tenant_id = nb.tenant_id
AND e.metric = nb.metric
AND date_trunc('hour', e.occurred_at) = nb.hour_start;
COMMIT;
Late events need no special path. An event that arrives at 15:20 for 13:00 lands in a new, small batch for the 13:00 hour with its own ID, and is added to the total. Stripe meters default to raw ingestion, where events for the same timestamp do not overwrite each other; the alternative, pre-aggregated ingestion, keeps only the most recent event per hourly or daily window (Stripe docs). The design above assumes raw.
Two edges need a written policy:
- Too old. Events older than 35 days are rejected by Stripe. Mark those batches and bill them as a manual invoice item, or absorb them, but never drop them silently.
- After the invoice. Decide a grace window after period end, such as 6 hours, before you treat a period as closed, and what happens to usage that arrives later. Stripe will not correct a finalized invoice for cancelled usage (Stripe docs).
How do you send meter events to Stripe without double billing?
Send each batch once, using the batch ID as Stripe's identifier, and only mark it sent after Stripe accepts it. If the worker crashes in between, the next run resends the same identifier and Stripe deduplicates it. Stripe recommends idempotency keys "to prevent reporting usage for each event more than one time because of latency or other issues" (Stripe docs). Treat that as protection against short-term retries; check the API reference for how long identifiers stay unique, and let your batch table be the long-term guard.
import Stripe from 'stripe'
import { Pool } from 'pg'
const stripe = new Stripe(process.env.STRIPE_SECRET_KEY as string)
const db = new Pool({ connectionString: process.env.DATABASE_URL })
const MAX_AGE_MS = 34 * 24 * 60 * 60 * 1000 // Stripe accepts 35 days; keep a margin
export async function sendPendingBatches(): Promise<void> {
const { rows } = await db.query(
'SELECT b.id, b.metric, b.hour_start, b.total, t.stripe_customer_id ' +
'FROM meter_batches b JOIN tenants t ON t.id = b.tenant_id ' +
"WHERE b.status = 'pending' ORDER BY b.hour_start LIMIT 500",
)
// Sequential on purpose: Stripe allows one concurrent call per customer per meter.
for (const b of rows) {
const hourStart = new Date(b.hour_start)
if (Date.now() - hourStart.getTime() > MAX_AGE_MS) {
await db.query("UPDATE meter_batches SET status = 'too_old' WHERE id = $1", [b.id])
continue // goes to a manual invoice-item queue, never silently dropped
}
try {
await stripe.billing.meterEvents.create({
event_name: b.metric,
identifier: b.id, // same batch, same identifier: a resend is deduplicated
timestamp: Math.floor(hourStart.getTime() / 1000),
payload: { stripe_customer_id: b.stripe_customer_id, value: String(b.total) },
})
await db.query("UPDATE meter_batches SET status = 'sent', sent_at = now() WHERE id = $1", [b.id])
} catch (err) {
if (err instanceof Stripe.errors.StripeRateLimitError) return // back off; next run retries
await db.query(
'UPDATE meter_batches SET attempts = attempts + 1, last_error = $2, ' +
"status = CASE WHEN attempts + 1 >= 5 THEN 'failed' ELSE 'pending' END WHERE id = $1",
[b.id, String(err)],
)
}
}
}
Stripe accepting the call is not the end. It processes meter events asynchronously and reports problems later as v1.billing.meter.error_report_triggered or v1.billing.meter.no_meter_found events, with codes such as timestamp_too_far_in_past or meter_event_customer_not_found (Stripe docs). Subscribe to both, alert on them, and fix and resend. Events sent in the last 24 hours can be cancelled with a meter event adjustment, and Stripe also lets you correct usage by recording a negative quantity (Stripe docs). If volume grows past the standard limit, Stripe's API v2 meter event streams accept up to 10,000 events per second in live mode (Stripe docs).
Test the sender like payment code: replay the same batch, kill the worker between the call and the update, and feed it an event from 36 days ago. The patterns are in testing payments and webhooks.
How do invoice previews, credits and prepaid plans work?
Previews. Customers will ask "what will I pay this month?" mid-cycle. Stripe can return a preview invoice, but meter totals on upcoming invoices "might not immediately reflect recently received meter events". Show live usage from your own table, priced with your own copy of the rate card, and label the Stripe preview as an estimate.
Credits and prepaid. Stripe credit grants cover prepaid and promotional credit. The rules that shape your design (Stripe billing credits):
- Credits apply only to subscription items on metered prices that report through Meters, not to licensed (flat or seat) prices or one-off invoices.
- They apply after discounts but before taxes, and only when the invoice finalizes, so credits shown on a preview can change.
- When several grants apply, priority wins, then earliest expiry, then promotional before paid.
- A customer can hold up to 100 unused grants, and credits cannot be offered as gift cards or general stored value.
- Voiding an invoice reinstates applied credit; issuing a credit note does not.
Mirror every grant and burn-down in your own ledger table so the in-app balance matches the invoice. For credit-heavy models, this is one of the reasons Stripe gives for choosing Metronome.
How do you enforce plan limits and send overage alerts?
Enforce limits in your own database, at the moment of use; alert from both your system and the billing system.
- Hard limits (free tier, prepaid exhausted): check usage this period against the plan allowance before doing the work. A sum over the indexed usage table, or a running counter updated in the same transaction, is fast enough for most products.
- Soft limits (overage allowed): notify at 80% and 100% of the included allowance, once per threshold per period, with the record of each alert stored so it is not sent twice.
- Stripe usage alerts can notify you when a customer crosses a meter threshold, with up to 25 alerts per meter and customer; billing thresholds can trigger an invoice early, but do not apply to trials and may invoice slightly above the threshold (Stripe usage monitoring).
Entitlements, the "may this tenant do this?" answer, should live in your database and be updated from webhooks, as covered in the multi-tenant SaaS guide. If those webhooks are arriving but plans are not changing, see Stripe webhook returns 200 but the subscription is not updated.
How long does it take to build usage-based billing yourself?
For a developer who knows Postgres and Stripe: about 3–5 days for one metric end to end (events, batching, sender, error webhook), and a further 2–3 weeks for credits, live usage pages, limits, alerts, reconciliation and the tests that prove a retry never double-bills. Those are our engineering estimates, not vendor figures.
The main risk of doing it yourself is silent mis-billing. Double-counting gets noticed fast because customers complain. Under-billing, from dropped late events, a dead worker or rejected events nobody watches, can run for months, and Stripe will not issue a corrected invoice for cancelled usage once an invoice is finalized. A nightly job that compares your totals with Stripe's per customer and period is low-cost insurance.
Why RAITHub for this
- Billing is part of the SaaS service. RAITHub's SaaS development scope includes Stripe subscriptions, metering, invoices and dunning, with entitlements in your database and idempotent webhooks.
- Metering on a live platform. The BlockEstate case study lists tiered partner API plans with "Stripe billing with usage metering", on a multi-tenant listing platform that reached MVP in 6 weeks (BlockEstate).
- Payments across many providers. PadhAI, the AI tutoring platform RAITHub built, runs 9 payment gateways; PropDesk collects rent through Stripe and ships with 1,024 automated tests.
- Counting under load without extra infrastructure. This site's rate limiting runs on Postgres with no Redis, the same "count it in the database, atomically" discipline a usage table needs (rate limiting without Redis).
When you don't need us
- One metric, low volume, already on Stripe. A Dashboard-created meter and a careful scheduled script may be enough while you test the pricing.
- Usage pricing is your whole business, with enterprise contracts. Buy Metronome or Orb and spend engineering on clean events; we can help with that part, but you may only need a few days of it.
- You have not decided the pricing model. Settle that first; billing models explained is a good start.
- You need tax or revenue recognition advice. That is a job for your accountant. This guide is general information; confirm with your adviser.
How RAITHub would build this
- Scope: usage event schema and idempotent ingestion; hourly aggregation with a written late-event and closed-period policy; an idempotent sender to Stripe Billing meters or Metronome with error-event handling; live usage pages, plan limits and overage alerts; a nightly reconciliation report.
- Timeline: added to a new SaaS build, it fits inside the 4–6 week fixed-scope SaaS or MVP range. Retrofitting metering into an existing backend with several metrics, credits and reconciliation is backend work in the 6–12 week range. If billing is already broken, a code rescue takes 2–4 weeks.
- You receive: automated tests and CI, including replay, crash and late-event tests; handover docs and runbooks for failed batches and disputed invoices; full IP under NDA.
- Next step: a free 15-minute technical audit of your pricing and current billing code, then a written fixed quote. Book the audit, or see the SaaS industry page first.
Frequently asked questions
What is usage-based billing in SaaS?
Charging customers for what they consume, such as API calls, tokens, messages or storage, instead of or on top of a flat fee. It needs metering: recording every billable event, aggregating it per billing period and sending totals to a billing system that rates and invoices them.
Should I use Stripe Billing meters or Metronome?
For a new integration, Stripe's documentation recommends Metronome, now part of Stripe. Billing Meters remains fully supported, and Stripe suggests staying on it if you already bill through it or need full compatibility with Connect, Checkout, Adaptive Pricing or Workflows.
How do I stop usage from being billed twice?
Make each layer idempotent: a unique constraint on tenant plus idempotency key when you ingest, a stable batch ID used as the meter event identifier when you send, and a status column so a batch is marked sent only after the billing system accepts it.
What happens to usage events that arrive late?
In your own system, they join a new batch for the hour they happened in. Stripe accepts meter events timestamped up to 35 days in the past and 5 minutes in the future. Beyond that, or after an invoice is finalized, you need a written policy, such as billing them as a separate item or absorbing them.
Can I offer prepaid credits with Stripe usage billing?
Yes, through credit grants, which apply to metered prices reporting through Meters, after discounts and before tax, at invoice finalization. They cannot be used as gift cards or general stored value. For heavy credit and commit models, Stripe recommends Metronome.
How do I show customers their usage in real time?
From your own usage table, not from Stripe. Stripe processes meter events asynchronously and Billing Meters reconciles at invoice time, so upcoming-invoice totals can lag. Price the live view with your own copy of the rate card and label it an estimate.
How long does RAITHub take to build usage-based billing?
Inside a new SaaS build, it fits the 4–6 week fixed-scope range. Adding several metrics, credits and reconciliation to an existing backend is typically 6–12 weeks. Scope is fixed in writing after a free 15-minute audit.
Related posts
Ready to discuss your project?
Book a free 15-minute technical audit with our engineering team.