Back to BlogArchitecture & Engineering

Rate Limiting Without Redis on Serverless: Postgres and Memory

Rupak Amin

Founder & Lead Engineer, RAITHub

11 min read

You can rate limit a serverless app without Redis by keeping the counters in the Postgres database you already run, updated with one atomic INSERT ... ON CONFLICT statement, so every instance sees the same count. Add a per-instance memory fallback for database outages. Give high-volume, low-stakes endpoints a memory-only limit, so they do not add a database write to every request.

This is how the RAITHub website limits login, the contact form, the admin API and uploads, on serverless hosting with no Redis. The algorithms themselves, fixed window, sliding window and token bucket, are explained in API rate limiting explained. This post is the implementation: the table, which store each endpoint uses, cleanup, failure handling, cost on a database that scales to zero, and when you should use Redis or a platform firewall instead.

Why doesn't a normal in-memory rate limiter work on serverless?

Because each instance keeps its own counter. A limit of 5 per minute across 4 running instances is really 20, and instances come and go with traffic, taking their counts with them. A limit that must hold needs a store that every instance reads and writes.

The usual answer is Redis. If you already run Postgres, and your limits protect logins and forms rather than millions of API calls, the database can be that store with no new service to provision, secure or pay for.

Which store should each rate limit use?

Choose by what a wrong count would cost. A limit that protects accounts, inboxes or money must be shared across instances. A limit that only slows spam can live in memory. This is the full set on this site, from src/lib/rate-limit.ts:

What it protectsLimitStoreWhy that store
Login, per email and IP5 a minutePostgresPassword guessing must be counted globally
Login, per IP across all emails20 a minutePostgresStops one IP trying many accounts
Login, per email across all IPs50 per 15 minutesPostgresSlows distributed guessing without locking out the owner
Contact form5 a minutePostgresSpam that reaches a real inbox
Confirmation emails, per recipient3 a dayPostgresAnyone can type anyone's address into the form
Admin writes30 a minutePostgresCaps bulk changes through the admin API
Media uploads10 a minutePostgresEach upload costs money at the image host
Public API120 a minutePostgresModest volume; fairness between callers
Analytics beacons60 a minuteMemory onlyLow stakes, very high volume; must not wake the database

Eight of the nine limits share one table. The ninth is the exception that keeps the database bill small, explained below.

What does the rate limit table look like?

Three columns: a key, a count and the time the window resets. This is the Prisma model:

model RateLimit {
  key       String   @id
  count     Int      @default(1)
  resetAt   DateTime

  @@index([resetAt])
}

The key is the primary key, so the upsert below has a unique column to conflict on. The index on resetAt exists for cleanup, which deletes by expiry time. Keys are SHA-256 hashes of identifiers such as a login email plus IP, so every row is a fixed 64 characters and no raw email address is stored.

The increment is one statement. PostgreSQL documents that ON CONFLICT DO UPDATE guarantees an atomic insert-or-update outcome, even under concurrency (PostgreSQL: INSERT). Written as plain SQL:

INSERT INTO "RateLimit" ("key", "count", "resetAt")
VALUES ($1, 1, $2)                      -- $2 = now + window
ON CONFLICT ("key") DO UPDATE SET
  "count"   = CASE WHEN "RateLimit"."resetAt" <= $3 THEN 1           -- $3 = now
                   ELSE "RateLimit"."count" + 1 END,
  "resetAt" = CASE WHEN "RateLimit"."resetAt" <= $3 THEN $2
                   ELSE "RateLimit"."resetAt" END
RETURNING "count";

The returned count is post-increment, so the first request in a window gets 1. If it is above the limit, the request is refused with a 429. The window resets inside the same statement, so there is no read-then-write race at the boundary. The rate limiting explainer walks through that statement line by line.

How do you clean up expired rows without a cron job?

Delete them opportunistically: on roughly 1 in 100 limiter calls, remove every expired row. This project has no scheduler for this job, and does not need one.

// Inside check(), before the upsert
if (Math.random() < 0.01) await cleanupRateLimits()

export async function cleanupRateLimits(): Promise<void> {
  try {
    await prisma.rateLimit.deleteMany({ where: { resetAt: { lte: new Date() } } })
  } catch {
    // Best-effort only; never fail a request because cleanup failed.
  }
}

The table stays bounded by traffic: busy periods clean up more often, quiet ones barely grow. The in-memory fallback uses the same idea, sweeping expired entries only when the map passes 10,000 keys.

What we deliberately avoided is setInterval. The code comment records why: on serverless it "never fires reliably", and when it does fire it keeps the instance alive. Timers assume a long-running process, which a serverless function is not.

What happens to the rate limiter when the database is down?

It fails open to a per-instance memory counter and logs the failure once per instance. Visitors can still send the contact form during a database incident; limits just become per-instance for a while.

} catch (error) {
  // Logged once per instance, not per request: an outage would otherwise
  // flood the logs. The single log still surfaces a missing table.
  if (!fallbackLogged) {
    fallbackLogged = true
    console.error('[rate-limit] Database limiter failed; using per-instance memory fallback:', error)
  }
  return checkMemory(identifier, options)
}

Failing open is a choice, not a default. It suits limits whose job is to slow abuse. If a limit guards something expensive, such as a paid API call per request, failing closed and returning 429 may be the safer choice during an outage. Decide per limiter, and write the decision down next to the code.

Why not use the database limiter everywhere?

Because every check is a database write, and some endpoints are called on every page view. On a database that scales to zero, those writes keep it awake and billing.

This site uses Neon Postgres. Neon's documentation says compute "scales to zero after an inactive period of 5 minutes" and reactivates within a few hundred milliseconds when queried again (Neon: scale to zero). The first-party analytics beacon fires on every page view. With a database-backed limit, every page view would also be a limiter write, and the database would stay awake as long as anyone was browsing.

So the beacon uses a separate, memory-only limiter:

// src/lib/rate-limit.ts
export function memoryRateLimit(options: { limit: number; windowMs: number }) {
  return {
    async check(rawIdentifier: string) {
      return checkMemory('mem:' + storageKey(rawIdentifier), options)
    },
  }
}

// src/app/api/track/route.ts
// Memory-only: a DB-backed limiter would add a Neon write to every page view.
const trackLimiter = memoryRateLimit({ limit: 60, windowMs: 60_000 })

It under-counts across instances, which is fine for spam control on a beacon. The route also drops bots, admin pages and non-production deploys before touching the database at all. The general rule: put the cheap filters first, and only pay for a shared count where a wrong count would hurt.

The same thinking applies to reads: public pages are cached with ISR and unstable_cache, so readers and crawlers rarely reach the database either.

When should you use Redis or a platform firewall instead?

When volume is high, when latency per check matters, or when the limit should stop traffic before it reaches your code at all. The options compare like this:

OptionShared across instancesExtra serviceGood for
Postgres table (this site)YesNoneLogin, forms, admin writes, modest APIs
Memory onlyNoNoneBeacons and spam control where under-counting is acceptable
Managed Redis over HTTPYesYesHigh-volume APIs and per-request quotas
Platform firewall rulePer regionPart of the hostBlocking floods before your function runs

For managed Redis on serverless, look for an HTTP-based client, because short-lived functions pay a price for opening and holding TCP connections. Upstash, for example, describes its rate limit library as connectionless and HTTP-based, designed for serverless and edge (Upstash: Ratelimit overview).

A platform firewall works at a different layer. Vercel's WAF rate limiting offers a fixed window on all plans and token bucket on Enterprise, with windows from 10 seconds, and its docs note that counters are tracked per region, so traffic across regions can exceed a single region's limit (Vercel: WAF rate limiting). It is a good outer wall against floods. It cannot key on things only your code knows, such as "the email address typed into the login form", so application limits still matter.

How do you test a database-backed rate limiter?

Test the logic in unit tests, and trust the atomicity to Postgres. This site's tests drive both paths of the limiter with a mocked database.

  • The database path, with a fake that emulates the upsert: under the limit, over it, remaining counts, independent keys, window expiry, a limit of 1 and a limit of 1,000.
  • Shared counts across instances. Two limiter objects model two serverless workers, and the test asserts the limit holds across both.
  • The fallback, by making the database reject: it still limits and never throws.
  • Stored keys are always 64 hex characters and never contain an @, even for a 5,000-character email.
  • Client IP parsing takes the rightmost x-forwarded-for value, including IPv6 and empty segments.

A mock cannot prove concurrency. If your limiter protects something valuable, add an integration test against a real Postgres that fires, say, 50 checks in parallel against a limit of 10 and asserts exactly 10 succeed.

Why RAITHub for backend work like this?

Because this limiter runs in production on this site, with its trade-offs written in the code, and the same habits go into client backends: choose the store by the cost of a wrong count, keep cheap filters first, and test the failure path as well as the happy one.

  • Cost-aware design. The memory-only beacon limit came from asking what each check costs on a scale-to-zero database.
  • Tests on the rules. Limits, keys and fallbacks are covered by automated tests gated in CI.
  • Fixed scope. Our API and backend development work is quoted in writing after a free technical audit.

When you don't need us

  • Your host's firewall covers it. If all you need is "no more than N requests per IP", a platform rule may be enough.
  • You already run Redis. Use it. A mature Redis rate limit library is less code than this.
  • You need very high request volumes. Use a purpose-built store and an engineer who runs it at that scale daily.

For the security side, see SaaS security best practices. To talk through your limits, book the free 15-minute technical audit.

Based on this site's src/lib/rate-limit.ts, running on Vercel with Neon Postgres. Last reviewed: 29 September 2026.

Frequently asked questions

Can I rate limit on Vercel without Redis?

Yes. Store counters in a Postgres table and increment them with one atomic INSERT ... ON CONFLICT DO UPDATE, so every function instance shares the count. It suits logins, forms and modest APIs; high-volume APIs are better served by Redis.

Is Postgres fast enough for rate limiting?

For most websites and admin APIs, yes. Each check is one small write on an indexed primary key. It adds a database round trip per limited request, so it is a poor fit for limiting every request on a busy public API.

How do I clean up an expired rate limit table without cron?

Delete expired rows on a small random fraction of limiter calls, such as 1 in 100, with an index on the reset time. The table stays bounded by traffic and you need no scheduler.

Should a rate limiter fail open or closed?

It depends on what it protects. This site fails open for forms and logins, so a database outage does not block real users. Limits in front of expensive paid calls may be safer failing closed.

Why use a memory-only rate limiter at all?

For high-volume, low-stakes endpoints such as analytics beacons. It under-counts across instances, but it adds no database write per request, which matters on a database that scales to zero.

Does Vercel have built-in rate limiting?

Yes. Vercel's WAF offers rate limiting rules, fixed window on all plans and token bucket on Enterprise, with per-region counters. It stops floods before your code runs, but cannot key on application data such as a login email.

Rate limitingServerlessPostgreSQLVercelNext.jsRedis alternativeNeon

Ready to discuss your project?

Book a free 15-minute technical audit with our engineering team.