Back to BlogQuality & Testing

Usability Testing for AI-Built Apps: What AI Cannot Catch

Rupak Amin

Founder & Lead Engineer, RAITHub

7 min read

Usability testing is watching a real person try to finish a real task in your app, noting where they hesitate, misread or give up. AI cannot do it: it checks that a button works, not whether anyone can find it or read the label. In an AI-built app, where code and tests share one prompt, usability is the gap no automated check covers. You run it with five people.

If you would rather have this run for you, see how RAITHub would test this below.

What is usability testing, and why can't AI do it?

Usability testing asks a real person to complete a task, such as "sign up and send your first invoice", while you watch without helping. You learn where the design confuses them. AI testing checks something different: that the code behaves as written. A model that generated the app from a prompt tests against its own assumptions, so it confirms the invoice endpoint returns a 200; it cannot notice that a first-time user never finds the "New invoice" button, or that the empty dashboard looks broken. In the 2025 Stack Overflow survey, 66% of developers named "AI solutions that are almost right, but not quite" as a frustration (Stack Overflow Developer Survey 2025). "Almost right" is usually a usability failure, not a crash.

What does usability testing catch that automated tests miss?

Problem a human seesWhat the automated test reports
New user lands on a blank dashboard and cannot tell what to doDashboard renders: pass
The primary action is below the fold on a phoneButton exists in the DOM: pass
An error says "Error 422" with no plain-English fixValidation returns the expected status: pass
A real address with an apostrophe is rejectedThe one seeded address is accepted: pass
The empty state looks like a bug, so the user leavesNo data, no crash: pass
Two steps that could be one, so people drop outBoth steps return 200: pass

Every row is a pass for the suite and a loss for the business. Onboarding is where it hurts most, because that is where new users decide whether to stay. Even well-run self-serve products convert only a median 8% of free sign-ups to paid (Growth Unhinged, 2026 free-to-paid conversion report); a confusing first screen pushes that lower, and no automated test will tell you why.

How do I run a usability test on my AI-built app?

  1. Write the top tasks. Three to five real jobs, phrased as goals: "create an account and add your first product", not "click New".
  2. Find five people. Five is enough to find most issues in a round; they should resemble your users, not your friends who already get it. Never use the person who prompted the app.
  3. Give a task, then stay quiet. Ask them to think aloud. Do not guide. The silence is where you learn.
  4. Watch on a real phone too. Generated layouts often break on small screens, which is where many of your users are.
  5. Note the moment of friction, not the opinion. "Looked for the button for 20 seconds" beats "I liked it".
  6. Rank and fix, then test again. Fix the drop-offs first; re-run with fresh people.

Do this before anyone tests code, so you fix the design before writing tests against it. For the structured version a founder can hand off, see the non-technical founder's QA guide and why users abandon your app.

Code sample: a usability test script you can reuse

Keep it plain text so anyone can run a session.

Task 1: Create an account and reach the main screen.
  - Start timer. Say nothing.
  - Note: where they pause, what they click by mistake, what they say.
  - Stop when they say "done" or give up.

Task 2: Add your first item (product / invoice / listing).
  - Watch the empty state. Do they know what to do with no data?

Task 3: Pay / subscribe (test mode).
  - Use a declined card: 4000000000000002 (Stripe test card).
  - Do they understand the decline message? Can they retry?

After each task, ask one question only:
  "On a scale of very hard to very easy, how was that?"

The declined-card line matters: real payments fail in ways the happy-path demo never shows, and Stripe publishes test cards for exactly these cases (Stripe testing docs).

Buy, build or hire usability testing?

RouteChoose this when
Build it yourself: a task list and five peopleYou have a first version, a few willing testers and a day. This finds most of the big issues cheaply.
Buy an unmoderated testing platformYou want recorded sessions from many people fast, and you can read and act on the videos yourself.
Hire a human QA pass with usability includedYou are launching to real users or investors and want usability, device, payment and access-control testing to a plan, with a ranked report.

Doing it yourself takes about a day plus the sessions; the main risk is testing with people who already understand the product, which hides the real friction. A hired pass removes that bias. At market developer and tester rates it is a small fraction of a build budget (Arc's 2026 freelance developer rates), and the gain is quality you can verify, shown by test counts rather than adjectives: 1,024 tests on PropDesk, 530+ on Sundor Skin, 750+ on TheSkinProof, the founder's own venture rather than a client. No usability case study exists yet, and none is implied.

How RAITHub would test this

  • Task-based sessions: real people attempting your top journeys on desktop and real phones, with a tester watching for the moments of friction.
  • Onboarding and empty states: the first-run experience, the blank dashboard, the error copy, the dead ends a new account hits.
  • Functional and access checks alongside: payments in test mode, the second user, the leaky endpoint, so usability and correctness are fixed together.
  • Ranked report with fixes: every issue with steps, evidence and a suggested change, ordered by what loses users first.

Timeline: a fixed-scope launch audit with dates agreed up front, then an optional monthly QA plan. You receive: a ranked report with reproduction steps and fixes; any automated tests go into your repository, with full IP and an NDA. Next step: a free 15-minute audit call, then a written fixed quote. See the AI-built app testing page, pre-launch QA, or the wider QA as a Service offer.

Unsure if people can use your app? Ask for a launch audit quote.

Frequently asked questions

Can AI do usability testing?

No. AI can check that a feature works, but usability is about whether a real person can find and understand it. That needs a human watching someone attempt a real task, which no automated test or AI agent provides.

How many people do I need for a usability test?

Around five per round finds most of the significant problems. Run a small round, fix the top issues, then test again with fresh people rather than gathering a large group once.

When should I run usability testing on an AI-built app?

As soon as a core journey works end to end, and before you write detailed tests against the design. Fixing confusing flows early is far cheaper than after they are coded and tested in place.

Is usability testing different from functional QA?

Yes. Functional QA checks the app does the right thing; usability testing checks a person can get it to do the right thing. AI-built apps usually need both, since the code and its tests share the same blind spots.

Do I need special tools to start?

No. A written task list, five representative people, a real phone and a willingness to stay quiet while they struggle are enough for a first, useful round. Tools and recordings help later.

usability testingai app usabilityvibe codingonboardingQA as a serviceux testing

Ready to discuss your project?

Book a free 15-minute technical audit with our engineering team.