Back to BlogQuality & Testing

Who Tests an App Built with AI? Your Options in 2026

Rupak Amin

Founder & Lead Engineer, RAITHub

9 min read

Usually no one. An app built with Lovable, Bolt, Cursor or Replit is tested by whoever the founder arranges: the founder clicking through, the AI tool's own scans and generated tests, an AI testing tool, a freelance tester, a crowdtesting platform or a QA-as-a-service team. Tools check the code against itself. A human tester checks it against real users, devices, payments and attackers. Most launches need both.

If you would rather have your AI-built app tested for you, see how RAITHub would test this below.

Why does almost no one test an app built with AI?

Because building got fast and testing did not. AI coding is now normal: 84% of developers use or plan to use AI tools (Stack Overflow Developer Survey 2025), and Google's DORA research puts AI adoption among software professionals at 90% (Google, 2025 DORA report). The same sources show the doubt that comes with it. In DORA's survey, 30% trust AI output only "a little" or "not at all". In Stack Overflow's, 66% name "AI solutions that are almost right, but not quite" as a frustration.

Testing lags behind writing. In the same Stack Overflow survey, 44% of developers said they don't plan to use AI for testing code, against 29% for writing code. So the code arrives faster than anyone checks it, and in a solo founder's project the person who would check it is often the same person who prompted it.

The cost of skipping it is concrete. Veracode tested code from more than 100 language models and found that 45% of samples introduced an OWASP Top 10 vulnerability (Veracode 2025 GenAI Code Security Report). The angle of this whole series is simple: AI tools write code fast, but they don't check it against real users, real devices, real payments or real attackers. Someone has to.

Who can test an AI-built app, and what does each one catch?

Who tests itCatchesMisses
You, the founderObvious breakage on the happy path, on your own laptop and phoneWhat you assume works, other roles, other devices, attacks
The AI that built it (generated tests, built-in scans)Known configuration issues, simple unit-level bugsBusiness rules it was never told, and its own blind spots
AI testing tools and agentsBroken selectors, regressions on recorded flows, quick exploration of pagesWhether the behaviour is right for your users; judgement calls
A freelance testerExploratory bugs, device issues, a fresh pair of eyesDepends on the person; rarely security or automation as well
CrowdtestingMany devices and locales in a short burstDeep knowledge of your roles, data and money flows
A QA-as-a-service teamFunctional, device, security, accessibility and regression testing to a planNothing it is not scoped for, so the scope must be written down

Can't the AI tool just test its own work?

It can help, and you should let it. Lovable runs a security scan that checks for tables without row-level security, open access rules and known-vulnerable npm packages; its own documentation says the quick scan does not examine application-specific logic, such as weak authorisation checks in your code, and that you remain responsible for the app's security (Lovable security docs). GitHub's documentation for Copilot test generation says "you should always review the generated code" and that the generated tests "may not cover all scenarios" (GitHub Docs).

That is the core limit. A model that wrote the code from a prompt tests against the same prompt. If the prompt never said "a member must not see another company's invoices", neither the code nor the tests will check it. How to test AI-generated code walks through the checks that close this gap, and why AI-written tests pass and still miss bugs covers the tests themselves.

What do AI testing tools add?

Speed on the repetitive part. Playwright now ships test agents: a planner that "explores the app and produces a Markdown test plan", a generator that turns the plan into test files, and a healer that repairs failing tests, or skips a test "if the healer believes that functionality is broken" (Playwright test agents docs). That is useful work. It still needs a person to decide whether the plan tests the right things and whether a skipped test is a bug in your product. Can AI test your app? goes through what these tools catch and miss, and AI testing tools vs a human tester compares them side by side.

What does a human tester find that tools don't?

  • The second user. Log in as user B and open user A's record. Cross-user and cross-tenant leaks are the bug class behind the public Lovable incident, CVE-2025-48757. See the Supabase RLS fix.
  • The real phone. A layout that works in a desktop preview and breaks on a small Android screen, or a keyboard that covers the pay button.
  • The real payment. A declined card, a 3D Secure challenge, a webhook that arrives twice. Stripe publishes test cards for exactly these cases, such as 4000000000000002 for a generic decline (Stripe testing docs), but someone has to use them.
  • The confused user. The empty state, the back button, the double-click, the form left open overnight.
  • The attacker. A request sent straight to the API with "role": "admin" in the body.

Which option should you choose at each stage?

StageWho should test
Prototype with test data, no real usersYou, plus the builder's scan. Work through the vibe-coded app security checklist
Before the first paying usersA one-off launch audit by a human tester, covering access control, payments and phones
Releasing every week with AI editsAutomated regression tests in CI, plus a monthly human pass on new features
Selling to larger customersSecurity and accessibility testing to a written scope, with reports you can share

Buy, build or hire?

RouteChoose this when
Buy a tool: an AI testing tool, a scanner or your builder's own security scanYou want a fast first pass and regression checks on flows you have already defined. Accept that it won't know your business rules.
Build it yourself: write the role map and tests with your AI toolYou can spend 3 to 5 days, you know the stack, and you will review every generated test by breaking the code on purpose.
Hire a freelancer or crowdtestingYou need many devices or a fresh pair of eyes for one release, and you can write the test plan.
Hire a managed QA-as-a-service teamReal users or real money are coming, and you want functional, device and security testing done to a plan with a written report.

For hiring routes in detail, read hiring a tester for your AI-built app. For budgets, see what testing an AI-built app costs.

Why RAITHub for this

  • Built for this gap. RAITHub tests apps built with Lovable, Bolt, Cursor, Claude Code, Replit, v0 and similar tools, starting where AI-built apps usually leak: access control, secrets and payments.
  • Test-heavy by habit. RAITHub's own builds carry large automated suites: 1,024 tests on PropDesk and 530+ on Sundor Skin. TheSkinProof, the founder's own venture rather than a client, runs 750+. There is no published case study of an AI-built app test yet; these counts show how RAITHub tests its own work.
  • Plain reports. Every bug comes with steps to reproduce, evidence and a suggested fix, ranked by what would hurt users first.

When you don't need a tester yet

  • The app is a demo on test data, with no sign-ups, no payments and no personal data.
  • You need a certified penetration test or a compliance attestation. RAITHub's security testing is application-level testing against OWASP guidance, not a CREST- or PCI-certified pentest.
  • You want a tester placed under your own management. RAITHub does not offer staff augmentation.

How RAITHub would test this

  • Role and flow map: what each kind of user should be able to see and do, written from your product rather than from the code.
  • Access control and secrets first: two accounts per role, direct API calls, a search of the browser bundle for keys, and row-level security checks on every exposed table.
  • Money and real devices: payment success, decline and webhook cases, then the main journeys on real iOS and Android phones and the common browsers.
  • Exploratory pass: a human tester working through the app the way a confused or hostile user would.

What you buy: a fixed-price launch audit for your AI-built app, with fixed dates agreed before it starts, then an optional monthly QA plan if you keep shipping AI edits. You receive: a written report of the bugs found, ranked, with fixes; any automated tests written go into your repository, with full IP and an NDA. Next step: a free 15-minute audit call, then a written fixed quote. Read more on the AI-built app testing page or the wider QA as a service offer.

Launching soon? Ask for a launch audit quote for your AI-built app.

Frequently asked questions

Who should test an app built with AI?

Someone other than the tool that built it. Use the AI tool's scans and generated tests as a first pass, then have a human tester check access control, payments and real devices before real users arrive.

Is it enough to ask the AI to write tests?

No. AI-written tests check what the prompt described, and they often pass against broken code. Review them by breaking the code on purpose and seeing whether anything fails.

Do Lovable's or Bolt's built-in checks replace testing?

They are worth running, but they don't replace it. Lovable's documentation says its quick scan does not examine your app-specific logic, such as weak authorisation checks, and that security remains your responsibility.

When should an AI-built app get its first human test?

Before the first real user enters personal data or a card. A launch audit at that point is far cheaper than fixing a data leak or a failed payment flow after it.

Can a non-technical founder arrange testing?

Yes. Write down who your users are and what each should be able to do, share the app and a test account, and ask for a ranked written report with steps to reproduce each bug.

Does RAITHub only test apps it built?

No. RAITHub tests apps built with any AI tool or by any team, from the live app, a staging copy or the repository, and reports in plain English.

testing app built with AIAI-built app QAvibe codingLovableCursorQA as a service

Ready to discuss your project?

Book a free 15-minute technical audit with our engineering team.