Back to BlogArchitecture & Engineering

A Scheduler and Recurring Jobs for Your SaaS

Rupak Amin

Founder & Lead Engineer, RAITHub

8 min read

RAITHub ships and tests production software. See QA as a Service or talk to us.

A scheduler runs work on a clock: nightly digests, billing retries, reminders, data cleanups. Build it on jobs stored in your own database, claimed atomically so two workers never run the same job, made idempotent so a retry or double-fire is harmless, retried with backoff on failure, and able to catch up a run the server missed.

If you would rather have scheduled jobs built for you, see how RAITHub would build this below.

What does a SaaS actually need a scheduler for?

For work that is not triggered by a user request: things that must happen on time, every time, whether anyone is logged in or not. RAITHub built this on PropDesk, a property-management SaaS with 1,024 tests that runs 5 daily automation jobs for rent, late fees, reminders and lease events, where a missed run means a tenant is not reminded or a fee is not applied.

Job typeExampleCost of a missed or doubled run
BillingRetry failed charges, apply late feesLost revenue, or a customer charged twice
CommunicationDaily digests, reminder emailsSilence, or a flood of duplicate emails
LifecycleTrial expiry, lease renewal, cleanupExpired trials still active, stale data piling up
MaintenancePurge old exports, roll up metricsGrowing storage bills, slow queries

This is the engine behind several other features: onboarding nudges, notification digests and data retention all run on it. Those features are in building a SaaS notification system.

Where should the schedule live, and who triggers it?

Keep the schedule and the job state in your database, and let a trigger (a platform cron, or a small ticker) wake a worker that reads what is due. Do not put the business logic in the cron itself; the cron only says "now".

  • Platform cron trigger. Most hosts can hit an endpoint on a schedule. Treat that endpoint as untrusted: require a secret, and have it only enqueue due jobs, never do the work inline.
  • Jobs in the database. A table of scheduled jobs with their cron expression and next-run time is the source of truth, visible and testable, rather than logic scattered across cron config.
  • A worker claims and runs. The worker picks up due jobs, runs them, records the result and computes the next run.

Postgres can back the whole queue and schedule, so you may not need a separate broker; the trade-offs for serverless mirror those in rate limiting without Redis on serverless.

How do you stop two workers running the same job?

Claim the job with a single atomic update, using FOR UPDATE SKIP LOCKED, so each due job is taken by exactly one worker and others move on. This is the standard Postgres queue pattern, documented under locking in the PostgreSQL SELECT reference.

-- One worker claims one due job; concurrent workers skip it and take the next.
WITH claimed AS (
  SELECT id FROM scheduled_jobs
  WHERE status = 'due' AND run_at <= now()
  ORDER BY run_at
  FOR UPDATE SKIP LOCKED
  LIMIT 1
)
UPDATE scheduled_jobs j
SET status = 'running', claimed_at = now()
FROM claimed WHERE j.id = claimed.id
RETURNING j.*;

Because the claim is atomic, you can run many workers and never double-process. After the work, set the job back to due with its next run_at, or to failed for retry.

How do you make a job safe to re-run?

Every job must be idempotent, because it will be re-run: a retry after a crash, a worker that claimed it and died, a schedule that fired twice. Make the effect depend on state, not on the fact that the job ran.

// Apply late fees once per invoice per day, even if the job runs twice.
export async function applyLateFees(db: Tx, asOf: Date) {
  // The WHERE clause is the idempotency guard: an invoice already fee'd today is skipped.
  await db.query(
    "INSERT INTO late_fees (invoice_id, fee_date, amount) " +
    "SELECT i.id, $1::date, i.late_fee FROM invoices i " +
    "WHERE i.due_date < $1 AND i.status = 'overdue' " +
    "ON CONFLICT (invoice_id, fee_date) DO NOTHING",
    [asOf],
  )
}

The unique key on (invoice_id, fee_date) means running the job twice applies the fee once. Prefer this to "remember we ran": state-based idempotency survives bugs that a ran-flag does not. The same discipline runs through imports and webhooks, as in CSV import, export and bulk actions.

What about failures, retries and missed runs?

Plan for the server being down when a job was due, and for jobs that throw.

  • Retry with backoff. On failure, schedule the next attempt after a growing delay, up to a cap, then move the job to a dead-letter state an operator can see.
  • Catch up missed runs. If the trigger did not fire at 02:00, the next tick should still find the job due and run it, because "due" is based on run_at <= now(), not on an exact instant.
  • Bound the catch-up. Decide whether a job that missed three days runs once or three times; for digests, once; for billing, exactly the right number.
  • Observe it. Alert when a job has not succeeded within its expected window, so a silently stuck scheduler does not go unnoticed for a week.

Buy, build or hire?

OptionExamplesChoose this whenWatch out for
Managed job/queue serviceHosted background-job and scheduling platformsYou want retries, dashboards and scheduling without running infrastructureJobs run in their system; idempotency and your data model are still yours
Platform cron onlyA host's cron hitting an endpointYou have one or two simple, infrequent jobsNo retries, no claim, no catch-up; business logic ends up in the endpoint
Build it in on your databaseThe claim-and-run pattern aboveJobs are core, must be idempotent, and should live beside your dataYou own the claim, retries and catch-up; test the concurrency
Hire a team to build itRAITHub or another studioScheduled billing or reminders must never miss or double, and must be provenInsist the handover includes the double-fire and missed-run tests

Do-it-yourself estimate: 3–6 days for the job table, atomic claim, idempotent runs, backoff retries, catch-up and alerting, if a trigger and database are in place. The main risk is a job that is not idempotent, so a retry or double-fire charges twice or emails twice.

How RAITHub would build this

As part of a new backend or SaaS build, or added to an existing product, scoped in writing after the free audit.

  • Jobs in your database: schedules and state as data, claimed atomically with SKIP LOCKED so workers never collide.
  • Idempotent runs: each job keyed so a retry or double-fire is harmless, backed by state rather than a ran-flag.
  • Retries and catch-up: backoff with a dead-letter state, and missed runs that catch up within bounds you set.
  • Observability: alerts when a job has not succeeded in its window, so a stuck scheduler is caught fast.

Timeline: this is a scoped piece inside a new 4–6 week SaaS build, or within the 6–12 week backend range when added to an existing system, with the exact scope in the quote.

You receive: automated tests and CI for the concurrency, retries and idempotency, handover docs and runbooks, and full IP assigned to you under NDA.

Next step: a free 15-minute technical audit, then a written fixed quote. See the API and backend development service, or book the audit.

Frequently asked questions

Can I just use my host's cron for scheduled jobs?

For one or two simple jobs, yes. But platform cron gives you no retries, no protection against two runs, and no catch-up if the server was down. Use it only to wake a worker that reads jobs from your database and does the work there.

How do I stop two workers running the same scheduled job?

Claim the job with a single atomic update using FOR UPDATE SKIP LOCKED, so one worker takes it and others skip to the next. This standard Postgres queue pattern lets you run many workers without ever double-processing a job.

Why must a scheduled job be idempotent?

Because it will be re-run: after a crash, a dead worker, or a double-fire. Make the effect depend on state, keyed so a repeat is a no-op, so running a billing or email job twice does the work once rather than charging or emailing twice.

What happens if the server was down when a job was due?

If "due" is based on run_at being in the past rather than an exact instant, the next tick finds the job and runs it. Decide in advance whether a job that missed several periods runs once or once per missed period.

Do I need Redis or a message broker for this?

Not to start. PostgreSQL can hold the schedule and act as the queue with SELECT ... FOR UPDATE SKIP LOCKED. A dedicated broker helps at high throughput, but many SaaS products run their jobs entirely on their existing database.

SchedulerRecurring jobsCronBackground jobsIdempotencySaaS

Ready to discuss your project?

Book a free 15-minute technical audit with our engineering team.