Back to BlogQuality & Testing

The Bugs Users Find First in AI-Built Apps, and How to Find Them Before They Do

Rupak Amin

Founder & Lead Engineer, RAITHub

12 min read

Users of AI-built apps tend to hit the same bugs first: layouts that break on phones, sign-up emails that never arrive, dead ends for brand-new accounts, double submissions, payments that do not unlock access, and errors that leak raw messages. Attackers find a second set first: other users' data, open admin pages and exposed keys. Each one has a cheap test you can run before launch.

If you would rather have these hunted down for you, see how RAITHub would test this below.

This is not a "we tested N apps" study. RAITHub has no AI-built-app case study yet. It is a list of bug categories, each backed by published research, vendor documentation or a public incident, and ordered by who finds it first. The angle is simple: AI tools write code fast, but they do not check it against real users, real devices, real payments or real attackers. Somebody has to.

Why do AI-built apps share the same bugs?

Because the tools build what the prompt describes, and prompts describe the happy path. In the 2025 Stack Overflow survey, 66% of developers named "AI solutions that are almost right, but not quite" as a frustration with AI tools (Stack Overflow Developer Survey 2025). "Almost right" means the demo works on the builder's laptop with the builder's account. The bugs sit in everything the demo skipped: a second user, an empty account, a small screen, a slow network, a declined card, a person who presses the button twice.

Which bugs do users find first, at a glance?

#Bug categoryWho finds it firstTest that finds it firstTime to check
1Broken layout on phonesUsers, on day oneEvery journey on a real iPhone and Android phone1–2 hours
2Sign-up and reset emails missing or brokenUsers, at sign-upSign up with Gmail and Outlook addresses on production30 minutes
3Dead ends for a new, empty accountUsers, in the first minuteFresh account, every screen, no seeded data1 hour
4Double submissions and duplicatesUsers with slow connectionsDouble-click every submit; throttle the network1 hour
5Paid, but no access (or the reverse)Paying customersDeclined cards, closed tabs, duplicate webhooksHalf a day
6Other users' data visibleAttackers, then the pressTwo-account test and an anonymous API callHalf a day
7Admin or internal pages open to anyoneAttackers and scannersRequest every admin route logged out and as a normal user1–2 hours
8API keys exposed in the browserAttackers, then your billSearch the build output for secret key prefixes15 minutes
9Fast with ten rows, slow with real dataYour first busy customerLoad realistic data volumes; check query plansHalf a day
10Each fix breaks something elseUsers, after every updateAn automated regression suite in CI1–2 days to set up

1. Why does the app break on phones?

Most AI building happens in a desktop browser with a wide preview pane, and most traffic is not desktop: mobile was 58.99% of worldwide web traffic in September 2026 (StatCounter). The usual failures are full-height panels sized with 100vh, which on mobile can bleed behind the browser toolbar (web.dev on viewport units), fixed widths that cause sideways scrolling, and menus that rely on hover. The fixes and a Playwright overflow test are in why an AI-built app breaks on phones.

2. Why do sign-up and password emails go missing?

Two common causes. First, the links inside the email point at localhost or a preview URL because the auth provider's site URL was never changed; Supabase's redirect URLs guide explains how the site URL and allow list control where those links go. Second, the email never reaches the inbox: Gmail requires every sender to authenticate with SPF or DKIM, and bulk senders with SPF, DKIM and DMARC (Google sender guidelines). A user who cannot verify their email never becomes a user, and you never see the error.

Find it first: sign up on the production domain with a Gmail and an Outlook address, click every link in every email, and request a password reset twice.

3. Why do new users hit a blank screen?

Because the builder's own account is full of data. Generated screens are often designed around a populated list, so a brand-new account sees an endless spinner, an error from an empty array, or a dashboard with nothing to click. This is the "almost right" problem in its plainest form.

Find it first: create a fresh account and visit every screen before adding anything. Each empty state should say what the screen is for and offer the next step.

4. Why do records appear twice?

A button that stays enabled while a request is in flight lets an impatient user, or a slow mobile connection, submit twice. Stripe's own go-live checklist tells developers to test with duplicate data, for example by retrying the same request. The same applies to every form in your app.

Find it first: a short Playwright test per create form:

import { test, expect } from '@playwright/test'

test('double-clicking Create makes one project', async ({ page }) => {
  await page.goto('/projects/new')
  await page.getByLabel('Project name').fill('Double submit check')
  await page.getByRole('button', { name: 'Create' }).dblclick()
  await page.goto('/projects')
  await expect(page.getByText('Double submit check')).toHaveCount(1)
})

Disable the button while saving, and enforce uniqueness on the server or in the database, so the guard does not depend on the browser.

5. Why do customers pay and still get no access?

Because the app unlocks access on the success page, and the success page may never load. Stripe's documentation says you cannot rely on fulfilment from the checkout landing page alone, because a customer can pay and lose their connection before it loads (Stripe fulfilment guide). Lovable's Stripe documentation notes that its integration does not use webhooks by default (Lovable Stripe docs), so check which path your app actually uses. The reverse bug, access without payment, comes from webhooks that skip signature checks. Both are covered in testing payments in an AI-built app.

6. Why can users see other people's data?

This is the most serious category, and the most thoroughly documented. Escape scanned 5,600 apps built on vibe-coding platforms and reported more than 2,000 high-impact vulnerabilities and 175 cases of exposed personal data, including medical records and bank details; misconfigured row-level security headed its list of causes (Escape, October 2025). CVE-2025-48757 described the same pattern in generated apps. Users rarely notice it; the first person to find it is usually someone looking.

Find it first: two accounts, one record ID, one refused request; then an API call with only the public key and no session. The full method is in testing login, roles and data access, and the fix for Supabase in RLS disabled in public.

7. Why are admin pages open to anyone?

Wiz's research on vibe-coded apps found internal tools, admin dashboards and staging environments deployed without authentication, and login checks written in browser JavaScript, including a password hard-coded in a script file (Wiz Research, September 2025). Hiding a menu item is not access control. Neither is a check that runs only in the browser.

Find it first: list every admin route and API, then request each one logged out and as a normal user. Each must refuse on the server.

8. Why are API keys showing up in the browser?

Escape found more than 400 exposed secrets in its scan, and Wiz found third-party keys, including AI provider keys, hard-coded in client-side code. A model key in the browser lets anyone run up your bill; a database service key lets anyone read everything. See exposed API keys in a Lovable app for the search command and the fix.

9. Why is the app fast in the demo and slow for customers?

The demo has ten rows. The first real customer imports five thousand, and a list page that fetches everything, or a query with no index, slows to a crawl. Supabase ships database advisors that flag missing indexes and other performance issues; run them, then load realistic volumes and measure. Performance testing before launch covers load tests.

10. Why does every fix break something else?

Each prompt solves its problem locally and can rewrite files it was not asked to touch. With no regression tests, nobody notices until a user does. The extreme version is public: in July 2025 an AI agent on Replit deleted a live production database during a test project despite instructions not to make changes (Fortune). Keep production separate, keep backups, and put a CI gate in front of every change; regression testing after every deploy shows how.

What about accessibility bugs?

They are found first by users who cannot complete the journey, and they rarely report it; they leave. They are common everywhere, not only in AI-built apps: WebAIM's 2025 scan of one million home pages found detectable WCAG failures on 94.8% of them, led by low-contrast text and missing form labels (WebAIM Million 2025). Generated UI inherits whatever the component library does, so check keyboard use and labels on every form.

In what order should you hunt these bugs on a small budget?

  1. Categories 6, 8 and 7: data, keys and admin access. A leak cannot be undone.
  2. Category 5 if you charge money at launch.
  3. Categories 1, 2 and 3: the first-minute experience on a phone.
  4. Category 4 on every create form.
  5. Categories 9 and 10 in the first weeks after launch, before growth makes them expensive.

Done yourself, the first four steps take roughly 2–3 days for a developer who knows the stack. The risk is testing like the builder: with a full account, a laptop and good intentions. The 40-point launch checklist for an AI-built SaaS turns this list into pass-or-fail checks, and exploratory testing explains how testers find the bugs no list predicts.

Buy, build or hire?

OptionChoose this whenTrade-off
A tool or SaaS testing platform (scanners, Playwright, a device cloud)A developer can write and maintain the tests aboveCatches known patterns well; misses what nobody thought to script
Freelancers or crowdtestingYou want many people and devices on the app for a launchBroad coverage of categories 1–4; little depth on 5–8
An in-house QA hireYou ship weekly and the product is past validationStrong over time; slow to hire for a first launch
A managed QAaaS team or a one-off launch auditYou want all ten categories checked before users arriveAn outside dependency; insist that reports and tests are yours

What each option costs is covered in QA as a service pricing.

Why RAITHub for this

  • Testers who think like your users and like attackers. RAITHub's testing covers both sets of bugs: exploratory testing on real phones for categories 1–5, and application security testing against OWASP guidance for 6–8.
  • Experience with the stacks AI tools generate. Supabase, Postgres row-level security, Next.js and Stripe are everyday tools in RAITHub's own builds, such as Sundor Skin with 146 tables under row-level security and PropDesk with 1,024 tests and Stripe rent collection.
  • Honest about proof. There is no AI-built-app case study yet; the test counts above are from RAITHub's own builds.

When you don't need us

  • Your app is a single-user tool or a landing page with no logins or payments. Categories 1 and 8 are enough.
  • You have a tester or a second developer who did not build the app. Hand them the table above.
  • You need a certified penetration test for a compliance audit. RAITHub's security testing is not CREST- or PCI-certified and produces no attestation.

How RAITHub would test this

  • Scope: a free 15-minute call about your app, your users, your phones and your payment provider.
  • Hunt: a tester works through all ten categories on a production-like environment, with two-account and logged-out tests for data access, real iPhone and Android devices, and test-mode payments.
  • Report: every bug ranked by severity, with reproduction steps, screenshots or recordings and a suggested fix.
  • Prevent repeats: optional monthly QA that turns the worst findings into automated regression tests in your repository.
  • What you receive: the report, any tests, and a handover note. IP is yours and an NDA is standard.

The launch audit is fixed-price, quoted in writing after the call. See AI-built app testing and QA as a service, or request a launch audit.

Frequently asked questions

What are the most common bugs in AI-built apps?

Users first meet broken phone layouts, missing or broken sign-up emails, dead ends for new accounts, double submissions and payments that do not unlock access. Security research also shows data exposed by missing access rules, open admin pages and API keys in browser code.

Are AI-built apps buggier than hand-written ones?

Not always, but their bugs are concentrated where prompts are vague: access rules, failure paths and edge cases. Veracode found an OWASP Top 10 weakness in 45% of AI code-generation tasks it tested (Veracode 2025).

How do I find bugs before my users do?

Use the app the way a stranger would: a fresh account, a phone, a slow network, a declined card, a second user trying to read the first user's data. Then automate the checks you want to repeat after every change.

Can my AI tool find these bugs itself?

It can help, for example by reviewing code for missing checks. But it shares the assumptions it built with, and it cannot hold a real phone or receive a real email. Treat its review as a first pass, not a sign-off.

Which bug should I fix first?

Anything that exposes other people's data or a secret key, then admin access, then payments. Those harm users or cost money in ways a later release cannot undo. Layout and empty-state bugs come next.

Does RAITHub have data on how many AI-built apps have these bugs?

No. RAITHub has not published a study of AI-built apps. The figures in this post come from the cited research by Escape, Wiz, Veracode, WebAIM and others.

ai app bugsvibe coded app bugscommon bugs in ai generated appslovable app bugsqa for ai-built appsexploratory testing

Ready to discuss your project?

Book a free 15-minute technical audit with our engineering team.