Back to BlogAI Features & LLMs

Why AI Can't Test Usability (and a Human Still Has To)

Rupak Amin

Founder & Lead Engineer, RAITHub

10 min read

No. AI can generate tests, click through pages and flag crashes, but usability is whether a real person can understand and finish a task, and that judgement is not in the code. An AI test checks that the app does what the app does. A confusing form, a dead-looking empty screen or a checkout that fails on one browser all pass that check and still lose the user.

If you would rather have a person test your app for you, see how RAITHub would test this below.

Can AI do usability testing on my app?

It can help with the mechanical parts and not with the judgement. An AI agent can open your pages, submit forms and report errors. What it cannot do is sit in a real person's head and notice that the label is unclear, the next step is hidden, or the error message blames the user for something the app got wrong. Usability is about expectation versus reality, and the expectation lives in your users, not your repository.

The tools' own makers say this plainly. GitHub tells Copilot users to "always review the generated code" and warns that generated tests "may not cover all scenarios" (GitHub Docs). Developers agree: in the 2025 Stack Overflow survey, 46% said they distrust the accuracy of AI tools and only 3% highly trust it (Stack Overflow Developer Survey 2025). If a person has to check the code an AI writes, a person has to check the experience an AI builds.

Why can't AI judge whether a screen is usable?

A test is a comparison between two things: what the app does, and what it should do. AI can observe the first with great speed. The second — what a confused first-time user expects, what feels broken, what is worth the effort — is not written in the code, so a model trained on the code cannot read it there. Four gaps follow from that.

  • Intent. "A new user should reach their first result without help" is a goal no model can infer from the function that renders the page.
  • Confusion. A form that technically works can still ask for the wrong thing in the wrong order. The code runs; the person gives up.
  • Trust. A payment page with an odd layout, a missing lock or a vague error makes people abandon even when the transaction would have gone through.
  • The empty and error states. A screen with nothing on it yet often looks broken. An AI that built the happy path rarely tests how it feels before any data exists.

There is a deeper reason AI cannot grade its own usability. A Stanford study found that people using an AI assistant wrote less secure code and were more likely to believe it was secure (Perry et al., ACM CCS 2023). The same confidence applies to experience: the model that built the flow is the worst judge of whether the flow is confusing, because it already knows what every screen is for.

What does AI catch, and what does a human catch?

Question about the appAI testingHuman tester
Does the page crash or error?Strong: finds broken flows and dead links fastConfirms and reproduces
Does the code do what the code says?Strong: that is exactly what it checksA poor use of a person's time
Can a first-time user finish the task?Weak: no sense of expectationStrong: watches where people hesitate
Is the onboarding confusing?WeakStrong: notices the step that loses people
Does the empty state look broken?Rarely tested at allStrong: tests the app with no data yet
Does checkout feel trustworthy?Can confirm it completesStrong: judges hesitation and drop-off
Does it work on a cheap phone on a weak signal?Partial, through device cloudsStrong: holds the device, feels the lag
Does the error message help or blame?Weak: reads it as a string, not as a reader wouldStrong

The pattern is consistent: AI is strong once a check is defined, and silent where no one has defined what "good" means. The place usability lives is exactly the undefined place. For the same split applied to automated tests, see why AI-written tests pass and still miss bugs, and for the whole comparison, real-user testing vs AI testing.

What usability failures actually kill AI-built startups?

Not dramatic crashes, usually. Quiet ones. The failures that lose a young company its first users tend to be about trust, edge cases and the moments a model never considered.

  • Onboarding that confuses. The person who signs up, cannot see what to do next, and leaves. The code worked perfectly.
  • A checkout that fails on one browser. A happy path tested in one place, breaking in another, taking real money's worth of customers with it.
  • A form that rejects a real address. Validation written for the common case rejects a flat number, a non-Latin name or a long postcode, and the real customer cannot buy.
  • An empty state that looks broken. A dashboard with no data yet, showing nothing, reading as "this app is broken" to a first-time user.
  • A silent data error. The screen says "saved" and nothing was. This is the one AI-written tests miss most, because they assert what the code already does. Veracode found that 45% of AI-generated code samples introduced an OWASP Top 10 vulnerability, and that newer models were no better (Veracode 2025 GenAI Code Security Report).

None of these show up in a passing test suite, because a suite written from the same prompt as the app inherits the same blind spots. That is why AI cannot fully test an app it helped build.

How do you actually test usability without a big budget?

The core method is older than AI and still works: watch a real person use the app and say nothing. The classic finding in usability research is that you do not need many participants — a handful of people trying the main journey surfaces most of the serious problems, because the same obstacles trip almost everyone.

  1. Write down the task, not the steps. "Sign up and get your first result," not "click here, then here."
  2. Find two or three people who have never seen the app. Ask them to do the task.
  3. Stay silent. Every time you want to help, write down why — that is a usability bug.
  4. Test the empty and error states on purpose: a brand-new account, a declined card (Stripe's test cards), a form left blank.
  5. Do it on a real phone, not only your laptop.

Doing this yourself takes about half a day to a day for a small app. The main risk is that you built it, so you already know where everything is and cannot un-know it — which is the exact reason an independent tester finds more. For founders who do not code, the non-technical founder's guide to getting your app tested turns this into a plain-English checklist.

Buy, build or hire?

RouteChoose this when
Buy a tool: AI test agents, your builder's security scan, an automated usability analytics toolYou want fast regression coverage and a record of where users click. Tools show what happened, not why it confused people.
Build it yourself: watch a few real usersYou can spend a day, you will stay silent and write down every hesitation, and you accept you cannot un-know your own app.
Hire a freelance tester or crowdtestingYou need fresh eyes on real devices for one release and can write the task list.
Hire a managed QA-as-a-service teamYou want a person to define what correct and usable mean, test the edge and empty states, and leave automated regression tests you keep.

Why RAITHub for this

  • A human decides what usable means. RAITHub uses AI-assisted tools where they save time, and has a person judge the experience, the edge cases and the empty states that AI reads past.
  • Test suites at real scale. PropDesk runs 1,024 automated tests, Sundor Skin 530+, and this website 400+. TheSkinProof, the founder's own venture and not a client, runs 750+. There is no published AI-built app case study yet; those counts are RAITHub's own builds.
  • Tests you keep. Any automated tests land in your repository and CI, so the next AI edit is checked too.

When to use a tool instead

  • Your app is a prototype on test data with no payments and no personal data. Run your builder's scan and a quick click-through.
  • You already have reviewed usability notes and only need regression tests generated from them.
  • You need a certified penetration test or a legal accessibility sign-off. RAITHub does neither: its security testing is application-level against OWASP guidance, and its accessibility audits test against WCAG 2.2 AA without certifying legal compliance.

How RAITHub would test this

  • Define usable: the tasks each kind of user should finish, agreed with you, which also becomes the brief for any AI test agent.
  • Human where AI is weak: onboarding, empty and error states, checkout trust, and the main journeys on real iOS and Android devices; see manual testing.
  • AI where it is strong: generated regression tests for the agreed flows, reviewed by a person, running in your CI.
  • A ranked report: every usability and functional issue with steps, evidence and a suggested fix, in plain English.

Buy it as a fixed-price launch audit with dates agreed up front, then add a monthly QA plan if you keep shipping AI edits. You receive the report, any tests in your repository, full IP and an NDA. See AI-built app testing, pre-launch QA and QA as a service. The next step is a free 15-minute audit call, then a written fixed quote.

Want a human check before launch? Request a launch audit quote.

Frequently asked questions

Can AI do usability testing?

It can explore pages, submit forms and report errors, which helps. It cannot judge whether a real person understands the screen, trusts the checkout or gives up on the onboarding, because that comparison lives in the user, not the code.

What is the difference between functional testing and usability testing?

Functional testing asks whether the app does what it is built to do. Usability testing asks whether a real person can actually use it to finish their task. An app can pass the first and fail the second, which is how a tested app still loses its users.

Why do my AI-written tests pass while users still struggle?

Because tests written from the same prompt as the app check that the code does what the code does. They inherit the prompt's blind spots, so a confusing flow or a rejected real address passes the suite and still fails the person.

How many people do I need for usability testing?

Fewer than most founders expect. Watching two or three people who have never seen the app attempt the main task surfaces most of the serious problems, because the same obstacles trip nearly everyone. Add more rounds as the app grows.

Can RAITHub certify my app is accessible?

No. RAITHub audits against WCAG 2.2 AA and reports each issue with its fix, which covers the usability of accessibility, but it does not certify legal compliance such as ADA, EAA or Section 508. Automated rules alone catch only part of the issues, about 57% per the axe-core project.

Should I still use AI testing tools if I pay for usability testing?

Yes. AI tools make regression testing cheap to build and keep running. A human defines what usable means and tests the parts automation cannot judge; the tools then protect those decisions on every change.

can AI do usability testingusability testingAI testing limitsvibe codingQA as a serviceAI-built app

Ready to discuss your project?

Book a free 15-minute technical audit with our engineering team.