Founder & Lead Engineer, RAITHub
Vibe-coded apps break in predictable places, so you can hunt the bugs on purpose instead of waiting for customers to find them. Go after eight targets in order: data access between users, declined and double payments, empty and error states, mobile layout, odd inputs, slow networks, logout and password reset, and anything the app did silently. Most of the damaging bugs hide in those eight.
If you would rather have a tester hunt them for you, see how RAITHub would test this below.
"Vibe coding" means building by prompting an AI tool and keeping what looks right, rather than reading the code. It is fast, and it ships working demos. The problem is that a demo exercises the happy path, and customers do not. The bugs that survive to launch are the ones nobody prompted a test for, and an AI test suite written from the same prompt inherits exactly those gaps.
Why do AI-built apps pass their tests and still have bugs?
Because a test written by AI from your prompt checks that the code does what the code does. It asserts the behaviour it already sees, so the suite stays green while a real failure ships. In the 2025 Stack Overflow survey, 66% of developers named "AI solutions that are almost right, but not quite" as a top frustration (Stack Overflow Developer Survey 2025). And the failures are not only cosmetic: Veracode found 45% of AI-generated code samples introduced an OWASP Top 10 vulnerability, with newer models no better (Veracode 2025 GenAI Code Security Report). To find these, you test against what should happen, which only you know. The deeper version of this is in why AI-written tests pass and still miss bugs.
Where do vibe-coded apps break? The eight targets
| # | Target | What to try | It has failed if |
|---|---|---|---|
| 1 | Data access between users | Make two customer accounts. Add private data in A. Log in as B and try to find or open it, including by pasting A's private URL. | B sees anything of A's |
| 2 | Wrong role | Log in as a basic user and open the admin URL directly | Any admin screen or data appears |
| 3 | Payments | In test mode, pay with a decline card, then double-click "Pay" | Access granted on decline, or two charges |
| 4 | Empty states | Use a brand-new account with no data yet | The screen looks broken or crashes |
| 5 | Odd inputs | Blank fields, a very long name, emoji, an apostrophe, a non-Latin address | Unclear errors, saved rubbish, or a rejected real address |
| 6 | Mobile layout | Every main screen on a phone, then an older one if you can borrow it | Buttons hidden, text cut off, forms unusable |
| 7 | Slow network and logout | Weak signal; then log out and press the browser back button | Double submissions, or private pages reappear |
| 8 | Silent failures | Do a save, then reload and check it really saved | It said "saved" and nothing was |
Targets 1 and 2 matter most: if either fails, stop and fix it before anything else, and do not invite real users yet. For the complete list of what customers hit first, see the bugs users find first in AI-built apps.
What's the single most valuable bug to hunt?
The cross-user leak — one person reading another person's data. It does the most damage and is the easiest for a vibe-coded app to get wrong, because AI often secures the screen while leaving the underlying request open. That is how CVE-2025-48757, a public Lovable data exposure, happened: the access rules were missing (CVE-2025-48757 disclosure).
You can hunt it without a browser at all. If you know one other account's record ID, ask the server for it directly as the wrong user and see whether it answers. In Playwright, that is one short test:
import { test, expect } from '@playwright/test'
test('a customer cannot read another customer order', async ({ request }) => {
const res = await request.get('/api/orders/' + process.env.OTHER_CUSTOMER_ORDER_ID, {
headers: { Authorization: 'Bearer ' + process.env.CUSTOMER_TOKEN },
})
expect([403, 404]).toContain(res.status())
})
If that request returns 200 with the order, the app has a leak, however locked-down the screens look. Payments are the other place to be deliberate: Stripe documents a decline card and a card that forces authentication on every charge, 4000002760003184 (Stripe testing docs), and nothing tries them unless a plan says so. Deeper payment coverage is in testing payments in an AI-built app.
How do I report a bug so the AI can fix only that?
Write every bug the same way, whether you paste it into your AI tool or send it to a developer:
- What I did: the exact steps from logging in.
- What I expected, from your written rules.
- What happened, with a screenshot or recording.
- Where: which account, browser and device.
Then add one line for the tool: "Fix only this. Do not change anything else." After the fix, re-run that check and re-run targets 1 and 2, because an AI fix in one place often breaks another. The bug report template has the full shape, and regression testing for AI-edited code explains why re-checking after every AI edit is not optional.
How long does this bug hunt take?
Allow one to two days for a small app if you can borrow a second phone and a spare email address: half a day on the access and payment targets, the rest on mobile, inputs and the silent failures, plus time to get each fix made and verified. The main risk of doing it alone is distance — you built the app, so you use it the way you intended, and the bugs hide in the paths you would never take. Watching one person who has never seen it try the main journey finds more than another day of your own clicking. When you are ready to go live, run it against a launch checklist.
What can't I hunt myself?
- Security beyond the cross-user checks: whether a technical attacker can reach data through paths the screens never show.
- Load: an app that works for you can fail with a hundred people at once.
- Payments in depth: refunds, failed renewals, and webhooks arriving late or twice.
- Device breadth: the full range of phones, browsers and screen readers your customers use.
This is where a tester earns their fee, and where an AI tool cannot help, because it shares your blind spots. The comparison is in real-user testing vs AI testing.
Buy, build or hire?
| Route | Choose this when |
|---|---|
| Buy a tool: your builder's security scan, AI test agents, a dependency audit | You want a fast first pass over common mistakes and repeatable regression runs. Someone still reviews what it produces. |
| Build the hunt yourself with your AI assistant | You have a day or two, you will write the cross-user tests, and you will re-check by breaking the code on purpose. |
| Hire a freelance tester or crowdtesting | You need fresh eyes on real devices for one release and can supply the targets. |
| Hire a managed QA-as-a-service team | You want a person to hunt the access, money and edge-case bugs and leave automated regression tests you keep. |
Why RAITHub for this
- Hunts where AI can't. RAITHub uses AI-assisted tools for speed and has a person chase the cross-user, payment and edge-case bugs that a prompt-written suite walks past.
- Test suites at real scale. PropDesk runs 1,024 automated tests, Sundor Skin 530+ (including a suite that tries to read other buyers' data), and this site 400+. TheSkinProof, the founder's own venture, runs 750+. There is no AI-built app case study yet.
- Tests you keep. Any automated tests land in your repository and CI, so the next prompt is checked too.
When to use a tool instead
- Your app is a prototype on test data with no payments and no real users.
- You already have a reviewed rules list and only need regression tests generated from it.
- You need a certified pentest or a legal accessibility sign-off. RAITHub's security testing is application-level against OWASP guidance, and its accessibility audits test against WCAG 2.2 AA without certifying compliance.
How RAITHub would test this
- Agree the targets: a role and flow map for your app, which also briefs any AI test agent.
- Hunt the high-damage bugs: cross-user and cross-role access, payments and webhooks, empty and error states, and the main journeys on real iOS and Android devices; see manual testing.
- Leave automation behind: generated regression tests for the agreed flows, plus hand-written cross-user checks, reviewed by a person and running in your CI.
- A ranked report: every bug with steps, evidence and a suggested fix.
Buy it as a fixed-price launch audit with dates agreed up front, then add a monthly QA plan if you keep shipping AI edits. You receive the report, any tests in your repository, full IP and an NDA. See AI-built app testing, pre-launch QA and QA as a service. The next step is a free 15-minute audit call, then a written fixed quote.
Want a human check before launch? Request a launch audit quote.
Frequently asked questions
Why does my vibe-coded app have bugs if the tests pass?
Because AI-written tests check that the code does what the code does. They assert the behaviour already in the code, so a missing rule or an untested path passes the suite and still breaks for a real customer.
What is the most important bug to find before launch?
A cross-user data leak — one person reading another's data. It does the most damage and is the easiest for a vibe-coded app to get wrong, because AI often locks the screen while leaving the underlying request open.
Can I hunt these bugs without reading code?
Most of them, yes. The access, payment, mobile, empty-state and input targets are about behaviour and need only two accounts, a phone and a decline test card. Security beyond the cross-user checks and load testing need tools or experience.
How do I stop the AI from breaking other things when it fixes a bug?
Ask for one change at a time, bookmark the working version first, and after every fix re-run your most important checks, not just the one you fixed. AI edits frequently cause regressions elsewhere.
Is my builder's security scan enough to catch these?
It is a worthwhile first pass for common misconfigurations like missing database rules, but it does not know your rules about which user may see which data, so the cross-user checks still matter.
When should I stop hunting myself and hire a tester?
Before real people trust the app with money or personal data, or when every AI edit seems to break something else. An independent tester finds what you cannot, because they did not build it and do not share your blind spots.
Related posts
Ready to discuss your project?
Book a free 15-minute technical audit with our engineering team.