Founder & Lead Engineer, RAITHub
An AI-built app is fast in the demo because the demo has one user and twenty rows. With 100 real users it slows down for a short list of reasons: query loops in the browser, row-level security on unindexed columns, no connection pooling, lists with no pagination, and AI calls made inside the request. Seed realistic data, run a k6 load test, and fix the slowest query first.
If you would rather have the load test run and the bottleneck found for you, see how RAITHub would test this below.
AI building tools are very good at producing a working screen from a prompt. They are not asked, and cannot check, how that screen behaves when a hundred people open it at once against a year of data. Nobody tested that, because in most AI-built projects nobody's job is testing. This post covers the performance side of that gap. For the general method (traffic estimates, test types, budgets), read load and performance testing before launch; this one is about what is different in apps built with Lovable, Bolt, Replit, v0, Cursor or Claude Code.
Why is my AI-built app fast for me but slow for real users?
Three things change between your demo and launch week, and generated code is rarely written with any of them in mind:
- Data volume. Your test account has 20 orders. A real customer has 2,000, and the admin screen loads all customers' orders. A query that scans a whole table is instant on 20 rows and slow on 200,000.
- Concurrency. One person clicking is one request at a time. A hundred people produce bursts of parallel requests that compete for the same database connections.
- Distance and devices. You test on a fast laptop next to the server region. Users arrive on mid-range phones over mobile networks, where every extra round trip from the browser is felt.
The model optimised for "the screen shows the right data". It did not optimise for how many queries that took, because the prompt never said, and in the preview it did not matter.
What makes AI-generated apps slow under load?
These are the patterns worth checking first. None of them is specific to one tool; they come from how prompts are phrased and how previews are tested.
| Pattern | How it appears in AI-built code | Symptom under load | Typical fix |
|---|---|---|---|
| Query loops (N+1) | Fetch a list, then one more query per item from the browser to get its owner or count | A page makes 50 to 200 requests; slow on phones | One query with a join or a database view; one API call per screen |
| Unindexed row-level security | Policies filter on user_id or org_id with no index, and call auth.uid() per row | Reads get slower as tables grow, even for small accounts | Index every policy column; wrap the function in a select |
| No pagination | select('*') on a whole table, filtered in the browser | Large payloads, memory spikes, slow first paint | Server-side filtering, limits and cursor pagination |
| No connection pooling | Serverless functions open direct database connections | Errors appear suddenly at a certain number of users | Use the transaction-mode pooler with a small pool per function |
| AI calls in the request | The page waits for an LLM response before rendering, with no cache or queue | p95 of several seconds; provider rate-limit errors | Stream, cache repeated prompts, move long jobs to a queue |
| Realtime everywhere | Every component opens its own live subscription | Connection counts climb with each open tab | One subscription per screen, only where live data matters |
Supabase's own documentation names the row-level security problem directly: "an unindexed filter column turns a read into a sequential scan", and wrapping functions such as auth.uid() in a select lets Postgres cache the result per statement rather than call it on each row (Supabase row-level security docs). Both fixes are one line of SQL:
-- Index the column the policy filters on
create index if not exists orders_user_id_idx on public.orders (user_id);
-- Re-create the policy with auth.uid() wrapped in a select
drop policy if exists "Users read own orders" on public.orders;
create policy "Users read own orders" on public.orders
for select to authenticated
using ((select auth.uid()) = user_id);
Check the policy still blocks other users afterwards. A speed fix that opens a data leak is worse than a slow page; the Supabase RLS fix guide covers how to test that.
How many users can a small database plan actually handle?
Fewer concurrent connections than most people assume, which is why pooling matters. On Supabase, the Nano and Micro compute sizes allow 60 direct database connections and 200 pooler clients; the Small size allows 90 and 400 (Supabase compute and disk docs). Connections are not users: 100 users browsing rarely need 100 connections at once. But a serverless app that opens a fresh connection per function call can use up 60 quickly during a burst.
Supabase recommends transaction mode on port 6543 "for serverless and edge functions, which open many short-lived connections", and notes that transaction mode does not support prepared statements (Supabase connection docs). If an AI tool wired your server functions to the direct connection string, that one change can remove a whole class of launch-day errors. Other hosts have their own limits; check yours before the test, not after.
How do you load test an AI-built app with realistic users?
Four steps, in this order. Skipping the first one is the most common reason a load test passes and launch still fails.
1. Seed realistic data
Fill a staging copy with the volume you expect after six to twelve months: thousands of users, tens of thousands of rows in the busiest tables, and at least a few heavy accounts. Generated test data is fine. An empty database makes every query look fast.
2. Create real test users
Make 50 to 100 accounts with different roles and log each one in, so the test exercises the same row-level security and caching that real sessions do. One shared test user hides most of the problems above.
3. Script the journeys that matter
Not the home page. The dashboard after login, the main list with filters, the action that writes data, and any screen that calls an AI model. A minimal k6 script that ramps to 100 virtual users, each with its own token, and pauses between actions like a person would:
// load/real-users.js: k6 run -e BASE_URL=https://staging.example.com load/real-users.js
import http from 'k6/http'
import { check, sleep } from 'k6'
const TOKENS = JSON.parse(open('./tokens.json')) // one token per test user
export const options = {
stages: [
{ duration: '2m', target: 25 },
{ duration: '3m', target: 100 },
{ duration: '10m', target: 100 }, // hold at 100 users
{ duration: '1m', target: 0 },
],
thresholds: {
http_req_failed: ['rate<0.01'],
'http_req_duration{name:dashboard}': ['p(95)<800'],
'http_req_duration{name:list}': ['p(95)<500'],
},
}
export default function () {
const params = {
headers: { Authorization: 'Bearer ' + TOKENS[(__VU - 1) % TOKENS.length] },
}
const dash = http.get(__ENV.BASE_URL + '/api/dashboard', Object.assign({ tags: { name: 'dashboard' } }, params))
check(dash, { 'dashboard 200': (r) => r.status === 200 })
sleep(2 + Math.random() * 3) // think time
const list = http.get(__ENV.BASE_URL + '/api/projects?page=1', Object.assign({ tags: { name: 'list' } }, params))
check(list, { 'list 200': (r) => r.status === 200 })
sleep(2 + Math.random() * 3)
}
Top-level stages in k6 is shorthand for the ramping-VUs executor, where a variable number of virtual users "executes as many iterations as possible for a specified amount of time" (k6 ramping-vus docs). Thresholds turn the budgets into pass or fail (k6 thresholds docs). If the app reads from the database directly in the browser, script those calls instead of an API: copy them from the network tab of a real session.
4. Watch the database while it runs
The k6 summary says the app got slow. The database says why. On Postgres, pg_stat_statements lists the queries that cost the most time in total:
select calls, round(mean_exec_time) as avg_ms, round(total_exec_time) as total_ms, query
from pg_stat_statements
order by total_exec_time desc
limit 10;
The top entry is usually a query loop or a missing index. Fix it, rerun the test, and repeat. Two or three rounds is normal. In Supabase, the index advisor in the Query Performance report suggests indexes for a selected query (Supabase index_advisor docs).
What should pass before you invite real users?
| Check | Target (engineering default) |
|---|---|
| Error rate at 100 users (or 2–3 times your expected peak) | Under 1% |
| p95 for list and read endpoints | Under 500 ms |
| p95 for the main dashboard or write action | Under 800 ms to 1 s |
| Requests per screen load | A handful, not dozens |
| Database connections at peak | Well below the plan limit |
| AI-backed screens | Stream or show progress; no blank page while waiting |
| Page experience on a mid-range phone | Largest Contentful Paint within 2.5 s, INP of 200 ms or less (web.dev Web Vitals) |
The server numbers are sensible starting budgets, not a standard. A load test measures your servers; it does not measure the page in the browser, so check Core Web Vitals separately. If production is slow for reasons that are not about load, such as rendering or caching, see why a Next.js app is slow in production, and if it only breaks after deploy, AI app works locally but breaks in production.
How long does it take to load test an AI-built app yourself?
Budget 2–4 days if you can read the code and the database: half a day to seed data and create test users, half a day to write and validate the k6 script, then one to two days of test, fix, rerun. If the fixes mean restructuring query loops across many screens, add days. The main risk is a false pass: an empty staging database, one shared test user, or a smaller plan than production will all make the result look better than launch day will be. A second risk is asking the same AI tool to "make it faster" without a measurement; it may add caching that serves one user's data to another.
Buy, build or hire?
| Option | Choose this when | Watch out for |
|---|---|---|
| A tool or SaaS testing platform (hosted load generators, APM) | You can script journeys and read database metrics | The tool creates load and graphs; diagnosing the slow query is still your job |
| Freelancers or crowdtesting | You want a one-off script and a report | Little help fixing the cause in AI-generated code |
| An in-house QA hire | Performance is a permanent concern and you release weekly | Load testing, database tuning and front-end profiling are rarely one person's skills |
| A managed QAaaS team | You want the test designed, run and diagnosed, with fixes named per query | Agree on a production-like environment and data volume first |
Why RAITHub for this
- The people who run the test read the code and the database, so the report names the query and the fix, not just "p95 too high".
- Bottlenecks traced, not papered over: on this site, a database hosting cost fix and rate limiting without Redis came from tracing real bottlenecks rather than adding servers.
- Honest proof: RAITHub builds ship with large test suites (1,024 tests on PropDesk, 530+ on Sundor Skin). There is no published case study of an AI-built-app load test yet.
When you don't need us
- You expect a few dozen users in the first months. Add the indexes, turn on pooling and run the k6 script above once.
- The slowness is one page and you can see the query loop in the network tab. Fix it directly.
- You need testers placed under your own management. RAITHub does not offer staff augmentation.
How RAITHub would test this
Through AI app testing, part of QA as a service, RAITHub would:
- agree the target load from your launch plan, and seed a staging copy with realistic data and test users per role
- script the key journeys in k6, including AI-backed screens, with payments pointed at the provider's sandbox
- run smoke, load and spike tests while watching queries, connections and provider limits
- report each bottleneck with the query or file, the cause and a specific fix, re-test after fixes, and check that no fix weakened access control
- leave the scripts in your repository with a small performance check in CI; see QA test automation
Buy it as a fixed-price launch audit for your AI-built app (a written report of the bugs and bottlenecks found, ranked by impact, each with a fix), with optional monthly QA afterwards. You own the scripts and all IP; an NDA is standard. The AI-built SaaS launch checklist covers what else to check before launch.
Next step: book a free 15-minute call about a launch audit, then get a written fixed quote.
Frequently asked questions
Why does my Lovable or Bolt app get slow as data grows?
Usually because queries scan whole tables: row-level security policies filter on columns with no index, lists load every row, or the page runs one query per item. Seed realistic data in staging and you will see it before users do.
Is 100 users a lot for a web app?
Not for a well-built one. It is enough to expose missing indexes, query loops and connection limits in an app that has only ever had one user, which is why it is a useful first target.
Does upgrading the database plan fix a slow AI-built app?
Sometimes briefly. More CPU and connections hide a missing index or a query loop until data grows again. Find the slow query first; upgrade if the fixed app still needs it.
Can I load test against production?
Before launch, use a production-like staging copy with realistic data. Never load test live payments or real customer accounts.
Can I ask the AI tool to make the app faster?
Yes, with a measurement. Give it the slow query from pg_stat_statements or the network tab and ask for that fix. A vague "optimise this" can add caching that leaks data between users, so re-test access control afterwards.
Is k6 free?
k6 is open source and runs locally or in CI. Grafana also sells a hosted cloud service for large or distributed tests, which is optional (k6 docs).
Related posts
Ready to discuss your project?
Book a free 15-minute technical audit with our engineering team.