Founder & Lead Engineer, RAITHub
RAITHub ships and tests production software. See QA as a Service or talk to us.
An EdTech app that runs fine all term and crawls on exam day has a concurrency problem, not a traffic problem. Hundreds of students arrive in the same five minutes and hit the same few endpoints. Find the slowest query under load, cap the database connection pool before it exhausts, move grading and PDF work off the request, and load-test the real exam shape, not the average day.
If you would rather have RAITHub diagnose and fix it for you, see how below.
This post is for EdTech founders and engineering leads whose exam or quiz platform slows down or times out at exam or enrolment time. It is engineering guidance. RAITHub's education build is PadhAI, an AI tutoring platform; the scaling patterns here draw on that and on this site's own performance work, and are labelled as such.
Why is my EdTech app slow only during exams?
Because exam load is spiky and correlated. For most of the term, requests trickle in. At the start of an exam, every student loads the same paper, starts an attempt, and begins autosaving, all within a few minutes. The same database rows, the same connection pool and the same grading code all get hit at once. A system sized for the average hour has no headroom for that minute.
| Symptom you see | Usual cause under load | Where to look first |
|---|---|---|
| Pages time out at exam start | Connection pool exhausted; requests queue behind slow queries | Database pool size and the slowest query in the log |
| Autosave fails intermittently mid-exam | Write contention on one hot table, or lock waits | Lock waits and row contention on the answers table |
| Submit is slow or times out | Grading or ranking runs inside the submit request | Whatever the submit handler does synchronously |
| Everything is slow, CPU low | Waiting on the database, not computing | Time spent in queries vs application code |
| Results page melts after the exam | Every student refreshes an uncached, expensive query | Caching and read models for results |
The generic version of this diagnosis is in why your Next.js app is slow in production. This post is the exam-season case of it.
How do I diagnose the exam-day slowdown?
Measure under the load that breaks it, not on a quiet afternoon. Three signals tell you most of the story.
- The slowest queries, at the time it slows down. Turn on slow-query logging and look at the exam window, not the daily average. PostgreSQL's pg_stat_statements extension tracks total and mean execution time per query, which finds the one query that dominates.
- Connection pool saturation. If active connections sit at the pool maximum and requests wait, you have found the bottleneck. More connections is usually the wrong fix; see below.
- Where the request spends its time. If the handler waits 900ms on the database and computes for 10ms, optimise the query or the round trips, not the code.
Reproduce it before you change anything. Write a load test that mimics the exam: many attempt-starts in two minutes, autosaves every 20 seconds, then a burst of submits. Run it against a staging copy. Load and performance testing before launch covers shaping a realistic test.
What are the common fixes, in order?
Work from the lowest-effort, highest-leverage change to the most involved. Most exam-day failures are fixed by the first three.
| Fix | What it solves | Rough effort |
|---|---|---|
| Index the hot queries | A sequential scan on a growing table that was fast when the table was small | Hours |
| Cap and tune the connection pool | Too many connections thrashing the database at peak | Hours |
| Move grading and PDFs to a queue | Submit requests that do heavy work inline | 1–3 days |
| Cache exam metadata and results | The same paper and the same results read thousands of times | 1–2 days |
| Open a lobby and stagger starts | Every student starting in the same 60 seconds | 1–3 days |
Why does adding more database connections make it worse?
Because a database serves only so many queries in parallel before context switching and lock contention slow every query down. Past that point, more connections mean more work competing for the same CPU and disk, so throughput falls. A bounded pool that queues requests briefly is faster than an unbounded one that overwhelms the database. PgBouncer's documentation describes exactly this: it lets you "have hundreds or thousands of database connections" from the application side while keeping "a small number of actual PostgreSQL connections" (PgBouncer features).
On serverless platforms this is sharper, because each function instance can open its own connections. Set a small per-instance pool and put a pooler in front of the database. A minimal connection-pool bound in TypeScript:
import { Pool } from 'pg'
// On serverless, keep the per-instance pool small: many instances x a big pool
// exhausts the database. Let a pooler (PgBouncer, or the platform's) fan in.
export const pool = new Pool({
connectionString: process.env.DATABASE_URL,
max: Number(process.env.DB_POOL_MAX ?? 5),
idleTimeoutMillis: 10_000,
connectionTimeoutMillis: 3_000, // fail fast instead of hanging a student's request
})
export async function withConn<T>(fn: (c: import('pg').PoolClient) => Promise<T>): Promise<T> {
const client = await pool.connect()
try {
return await fn(client)
} finally {
client.release()
}
}
This site runs rate limiting without Redis and had its database cost brought down by right-sizing exactly these settings. The wider patterns are in rate limiting without Redis on serverless.
How do I keep submit and autosave fast at peak?
Make submit a cheap status change, and make autosave a small, indexed write. The expensive work happens afterwards, off the request path.
- Submit records that the attempt is done, nothing more. Grading, rank calculation and certificate PDFs go to a background queue. A student pressing submit should wait on one row update, not on marking 60 questions.
- Autosave writes one answer, not the whole paper. Keep the answers table lean and indexed on the attempt, so a write at exam peak touches one row.
- The timer lives on the server. A server-owned deadline means a slow page or a refresh never costs a student time, and auto-submit runs as a job. The full pattern, with code, is in online exam software: timers and anti-cheating that works.
- Serve results from a cached read model. After the exam, everyone refreshes results at once. Compute ranks once, cache them, and serve reads from the cache.
Background jobs and webhooks under load have their own failure modes; testing payments and webhooks covers testing async paths, which applies to grading queues too. Live video is a separate correlated-load problem, covered in scaling live classes and video for thousands of students.
Buy, build or fix: how should you handle exam-season load?
Decide by whether exams are your product. If you run occasional quizzes, a hosted tool absorbs the spike for you. If exams are what customers pay for, the scaling is yours to own, or to hand to someone who will.
| Option | Choose this when | Watch out for |
|---|---|---|
| Hosted exam tool | Low-stakes quizzes; a vendor's concurrency cap covers your class sizes | Caps on simultaneous test takers; limited control over the spike and your data |
| Fix your own stack | You have the engineering time and the slowdown is a handful of queries and the pool | The real risk is a timer or autosave bug that only shows under load |
| Bring in help to diagnose and load-proof | The next exam is close and you cannot afford another failure | Scope the audit tightly: find the bottleneck first, then fix |
How long does it take to fix this yourself?
If you are comfortable with PostgreSQL and your framework, finding and indexing the worst query and capping the pool is often a day or two. Moving grading to a queue and adding a results cache is usually another few days. The main risk of doing it yourself is shipping a change you cannot load-test before the next real exam, when you cannot reschedule a thousand students. Build the exam-shaped load test first, so every fix is proven against it. Regression testing on every deploy keeps those tests from rotting.
How RAITHub would fix this
Scope:
- A performance audit under exam-shaped load: the slowest queries, pool saturation and where each request spends its time.
- Indexing and query fixes for the hot paths, and a tuned, bounded connection pool with a pooler in front.
- Grading, ranking and document generation moved to a background queue, with submit reduced to a status change.
- A cached results read model and a lobby that staggers attempt starts.
- A repeatable load test shaped like your largest exam, kept in CI.
Timeline: a focused diagnose-and-stabilise pass is a 2–4 week code rescue engagement; deeper scaling work on an exam platform follows the SaaS development path. You receive: automated tests and the load test running in CI, handover docs and an exam-day runbook, and full IP under an NDA signed before detailed discussion. Production and student data stay in your own cloud account; development uses synthetic data. Next step: a free 15-minute technical audit, then a written fixed quote; RAITHub publishes no rates. See what RAITHub builds for education on the EdTech industry page, and book the free audit with your largest exam size, your current stack and when the next exam runs.
Frequently asked questions
Why is my EdTech app fast all term and slow on exam day?
Exam load is spiky and correlated: hundreds of students start the same exam in the same few minutes and hit the same queries, rows and connection pool at once. A system sized for the average hour has no headroom for that peak, so it queues and times out.
Will a bigger server fix an exam-day slowdown?
Sometimes a little, but usually not the root cause. Most exam-day failures come from a slow query, an exhausted connection pool, or heavy work done inside the submit request. Fix those first; a bigger server without them just delays the same wall.
Why does adding more database connections make things slower?
A database runs only so many queries in parallel before contention slows every query. Beyond that, more connections compete for the same CPU and locks, so throughput falls. A small bounded pool with a connection pooler in front is faster than a large unbounded one.
How do I load-test for exam season?
Mimic the real shape: many attempt starts in two minutes, autosaves every 20 seconds, then a burst of submits at the deadline, all against a staging copy. Testing the average day will not reveal the peak that breaks production.
How do I stop submit from timing out under load?
Make submit record only that the attempt is finished, then do grading, ranking and PDFs in a background queue. A student pressing submit should wait on a single row update, not on marking every question synchronously.
Can you fix this before our next exam?
Often yes, if the next exam is a couple of weeks out. A tight diagnose-and-stabilise pass finds the bottleneck, fixes the hot queries and pool, moves grading off the request, and proves it with an exam-shaped load test. The timing depends on your stack, agreed in the audit.
Related posts
Ready to discuss your project?
Book a free 15-minute technical audit with our engineering team.