Back to BlogArchitecture & Engineering

Plan-Based Rate Limiting and Quotas in a SaaS

Rupak Amin

Founder & Lead Engineer, RAITHub

8 min read

RAITHub ships and tests production software. See QA as a Service or talk to us.

Rate limits cap requests in a short window (say 100 per minute); quotas cap usage over a billing period (say 10,000 API calls a month). Both should vary by plan, be enforced per tenant, and be counted atomically so two requests can never both take the last unit. Rate limits protect your system; quotas enforce what a customer paid for. Mixing them up is the common mistake.

If you would rather have rate limiting and quotas built for you, see how RAITHub would build this below.

What is the difference between a rate limit and a quota?

A rate limit answers "too fast?"; a quota answers "too much this month?". They have different windows, different purposes and different responses, and a SaaS needs both.

Rate limitQuota
WindowSeconds to minutesA billing period
PurposeProtect your system from spikes and abuseEnforce the plan the customer bought
Response when exceeded429, retry after a few seconds403 or an upgrade prompt; retrying will not help
ResetsContinuously, as the window slidesAt the start of the next period
Tied toThe endpoint and the tenantThe plan and the metered feature

Quotas are really entitlements with a counter, so the model and the atomic consume pattern overlap with SaaS entitlements, plans and limits. This post focuses on building both as plan-aware, per-tenant controls.

Which rate-limiting algorithm should you use?

A fixed window is simplest and lets a burst through at the boundary; a sliding window or token bucket is smoother. For most SaaS, a sliding-window counter is the right default: close enough to accurate, cheap to compute.

  • Fixed window. Count per clock minute. Easy, but a user can send the full limit at 10:00:59 and again at 10:01:00, double the intended rate for a moment.
  • Sliding window. Weight the previous window by how far into the current one you are. Smooth, and one counter per window is enough.
  • Token bucket. Tokens refill at a steady rate; each request spends one. Good when you want to allow short bursts but a steady average.

You do not need Redis to do this on serverless; a Postgres-backed or in-memory limiter works, which is how this site's own limiter is built. The trade-offs are in rate limiting without Redis on serverless.

How do you make the limit vary by plan, per tenant?

Look up the tenant's limit from its plan, then count against a key that includes the tenant and the window. The limit is data; the algorithm is code.

-- Atomically count a request in the current window, returning the new count.
-- period_key is e.g. 'minute:2026-10-11T09:03' for a rate limit,
-- or 'month:2026-10' for a quota.
INSERT INTO usage_counters (tenant_id, meter, period_key, used)
VALUES ($1, $2, $3, 1)
ON CONFLICT (tenant_id, meter, period_key)
DO UPDATE SET used = usage_counters.used + 1
RETURNING used;
export async function enforce(db: Tx, tenantId: string, meter: string, periodKey: string, limit: number) {
  const { used } = await db.one(/* the INSERT ... RETURNING above */, [tenantId, meter, periodKey])
  if (used > limit) {
    const err = new Error('rate_limited') as Error & { status: number; limit: number }
    err.status = 429
    err.limit = limit
    throw err
  }
}

Because the database does the increment-and-return in one statement, two simultaneous requests cannot both read "99" and both write "100". For a hard quota you want to block before the action, use the conditional variant that refuses when the increment would exceed the limit, exactly as in the entitlements counter.

What should the 429 response tell the caller?

Enough to back off correctly. A bare 429 makes clients retry blindly and makes the problem worse.

  • Retry-After, in seconds, so a well-behaved client waits exactly long enough.
  • The limit, remaining and reset, commonly as RateLimit-Limit, RateLimit-Remaining and RateLimit-Reset headers, so clients can self-throttle before they hit the wall.
  • A machine-readable body with a code, so the caller can tell a rate limit (slow down) from a quota (upgrade), which need different handling.

If you expose a public API, document these limits and headers; the broader design is in designing a public API for your SaaS.

Where should quotas be enforced, and what happens at the limit?

Enforce quotas on the server, at the action that consumes the unit, and decide the overage policy before launch.

  • Soft limit (warn). Notify the admin at 80% and 100% but keep serving, then reconcile. Good for goodwill; needs a billing or follow-up plan.
  • Hard limit (block). Refuse the action at 100% with a clear upgrade path. Simple and predictable.
  • Overage billing. Keep serving and bill the excess, which ties into metered billing; see usage-based billing for SaaS.

Whichever you choose, count usage even for unlimited plans, because that data tells you where to set next year's limits. The period reset and any overage reconciliation run as scheduled jobs; see a scheduler and recurring jobs for your SaaS.

Buy, build or hire?

OptionExamplesChoose this whenWatch out for
API gateway or edge limiterA gateway or CDN rate-limit featureYou want coarse per-IP or per-key limits at the edge before traffic hits your appHard to make per-tenant and plan-aware; quotas still need your app
Metering or billing platformUsage-metering SaaS tied to billingQuotas must line up exactly with invoices and overage billingRate limiting is still yours; another service on the path
Build it inThe counter and plan lookup aboveLimits vary by plan, must be per tenant, and live in your own dataYou own the algorithm choice and the headers; test the races
Hire a team to build itRAITHub or another studioRate limits, quotas and billing must agree and be proven under loadInsist the handover includes the concurrency tests, not just the limiter

Do-it-yourself estimate: 3–6 days for plan-aware limits, atomic counters for both windows, the 429 headers and a quota policy, if a usage-counter table exists. The main risk is a non-atomic read-then-write counter that lets concurrent requests slip past the limit.

How RAITHub would build this

As part of a new backend or SaaS build, or added to an existing product, scoped in writing after the free audit.

  • Plan-aware limits: rate limits and quotas read from the tenant's plan, enforced per tenant on the server.
  • Atomic counters: increment-and-check in one statement for both the short window and the billing period, safe under concurrency.
  • Clear responses: 429 with Retry-After and limit headers, and a machine-readable code that distinguishes rate limit from quota.
  • Overage policy: soft, hard or billed, implemented and tested, with usage recorded even on unlimited plans.

Timeline: this is a short scoped piece inside a new 4–6 week SaaS build, or within the 6–12 week backend range when added to an existing API, with the exact scope in the quote.

You receive: automated tests and CI for the concurrency and limit boundaries, handover docs, and full IP assigned to you under NDA.

Next step: a free 15-minute technical audit, then a written fixed quote. See the API and backend development service, or book the audit.

Frequently asked questions

What is the difference between rate limiting and a quota?

A rate limit caps requests in a short window to protect your system; a quota caps usage over a billing period to enforce the plan a customer paid for. Hitting a rate limit means slow down; hitting a quota means upgrade, and retrying will not help.

Do I need Redis to rate limit a SaaS?

No. A Postgres-backed or in-memory limiter works, including on serverless, which is how this site's own limiter is built. Redis helps at high scale or across many instances, but it is not required to start.

How do I keep two requests from both using the last unit?

Count with a single database statement that increments and returns the new total, or that refuses when the increment would exceed the limit. A read in application code followed by a separate write has a race that lets both requests through.

What headers should a 429 response include?

Retry-After in seconds, and limit, remaining and reset headers so clients can self-throttle. Include a machine-readable code in the body so the caller can tell a rate limit from a quota and handle each correctly.

Should limits be the same for every customer?

No. Limits should vary by plan and be enforced per tenant, so a higher plan genuinely gets more. Keep the limit as data read from the plan, and the enforcing algorithm as code, so pricing changes do not require code changes.

Rate limitingQuotasPlan limitsThrottling429SaaS

Ready to discuss your project?

Book a free 15-minute technical audit with our engineering team.