# How to test a vibe-coded app with real users

> Vibe-coded apps work for the person who prompted them. Five steps to test yours with real users and turn what they struggle with into fixes for your agent.

Page: https://hallwaytest.ai/learn/test-vibe-coded-app-with-real-users
Author: Kseniia Radova, Usability tester
Published: September 26, 2026. Last updated: September 26, 2026

**In short:** a vibe-coded app has usually been used by exactly one person, the one who prompted it. To find out whether it works for anyone else, give it to three to five people who have never seen it, one realistic task each, ask them to think aloud and watch without helping. Fix what they trip over, then repeat with new people. You can run the whole loop yourself or from your coding agent.

## Why vibe-coded apps need real users

Andrej Karpathy coined the term in February 2025: "There's a new kind of coding I call 'vibe coding', where you fully give in to the vibes, embrace exponentials, and forget that the code even exists." By November, Collins had made vibe coding its [Word of the Year](https://blog.collinsdictionary.com/language-lovers/collins-word-of-the-year-2025-ai-meets-authenticity-as-society-shifts/), defined as "the use of artificial intelligence prompted by natural language to write computer code."

The method changes who checks the work. When nobody reads the diffs, the only review a feature gets is its author clicking through it, and the author is the one person who cannot get confused by it. They know what every button does because they asked for it.

The gap shows up after launch. Nielsen Norman Group describes the pattern in its 2026 piece on [UX debt from AI](https://www.nngroup.com/articles/ai-ux-debt/): "AI lets teams build faster than UX can evaluate," and the result "may work well enough to demo or ship, but user confusion and degraded trust surface later." Developers see the same thing in code. In the [2025 Stack Overflow survey](https://survey.stackoverflow.co/2025/ai), the most common frustration with AI tools, named by 66%, was "AI solutions that are almost right, but not quite."

A usability test is the cheapest way to find the "not quite" before your customers do.

## Five steps to your first test

| Step | What to do | What you end up with |
| --- | --- | --- |
| 1. Pick one flow | Choose the flow that decides whether the app succeeds | The flow and its finish line |
| 2. Write the task | Describe a real situation, not the clicks | One scenario without clues |
| 3. Find testers | Three to five people who have never seen the app | A short list of sessions |
| 4. Watch them think aloud | Record screen and voice, do not help | The moments where people got stuck, in their words |
| 5. Hand it to your agent | Give evidence, not conclusions | Two or three fixes, then the next round |

## 1. Pick one flow and a finish line

Test the flow that decides whether the app succeeds, not the whole app. For most products that is one of three: signing up and reaching the first useful screen, doing the core action for the first time, or paying.

Write down what "done" looks like, for example "the tester has created a project and invited a teammate". Without a finish line, you cannot tell a success from a tester who gave up politely.

## 2. Write the task as a situation, not an instruction

How you phrase the task decides what you learn. Nielsen Norman Group's guidance on [task scenarios](https://www.nngroup.com/articles/task-scenarios-usability-testing/) comes down to three rules:

- **Make it realistic.** Give the tester a reason to be there that a real customer would have.
- **Make it actionable.** "It's best to ask the users to do the action, rather than asking them how they would do it."
- **Avoid clues.** [Remove any words that appear in your interface](https://www.nngroup.com/articles/better-usability-tasks/), or the tester will simply hunt for the matching button.

Compare two versions of the same task:

- **Instruction:** "Click New Project, add three tasks and share it."
- **Scenario:** "Your team is planning an offsite next month. Set things up so everyone can see what still needs to be booked."

The first one tests whether people can read. The second tests whether your app makes sense.

## 3. Find people who have never seen it

You need fewer than you think. Jakob Nielsen's research found that [five users uncover about 85% of the usability problems](https://www.nngroup.com/articles/why-you-only-need-to-test-with-5-users/) in a design, and his advice is that "the best results come from testing no more than 5 users and running as many small tests as you can afford." Steve Krug's [Rocket Surgery Made Easy](http://sensible.com/downloads/rsme-toc.pdf) scales it down further: three people, one morning a month.

For the basics, almost anyone who is not you will do. Krug's advice is to "recruit loosely and grade on a curve". You need people from your target audience only when the task depends on domain knowledge, such as an accounting feature or a clinical workflow.

Friends are available but polite, and they rarely give up or say "this makes no sense". Strangers are more useful and usually expect to be paid. The [User Interviews incentives report](https://www.userinterviews.com/blog/research-incentives-report), based on nearly 20,000 projects, puts the typical rate at $80 an hour for a remote consumer session and recommends $100 an hour as a baseline for professionals.

## 4. Watch them think aloud, and do not help

Ask each tester to say what they are thinking as they go: what they are looking for, what they expect to happen, what surprises them. Jakob Nielsen calls [thinking aloud](https://www.nngroup.com/articles/thinking-aloud-the-1-usability-tool/) "the single most valuable usability engineering method", and it is the difference between knowing that someone clicked the wrong button and knowing why.

While they work:

- Record the screen and their voice, so you can go back to the exact moment later.
- Do not explain, hint or apologise. If they ask what to click, ask what they would do if you were not there.
- Write down where they hesitate, go back, re-read or give up, with their exact words.
- Stop when they reach the finish line or clearly give up. Both are results.

As NN/g puts it, "the point of usability testing is to see what users do, not to hear what they would do." Opinions about colours and fonts are noise. The moments where behaviour breaks down are the findings.

## 5. Hand the findings to your coding agent

A vibe-coded app gets fixed the way it was built: by prompting. Give the agent evidence rather than conclusions, one friction point at a time:

```text
Tester 2, 04:12: on the plan step, the tester scrolled past the
Continue button twice and said "I don't know how to get to the next page".
Tester 4 did the same at 03:50. Make the next step obvious on a 13-inch
laptop screen without changing the plan cards.
```

Fix the two or three problems that stopped people, not everything anyone mentioned. Then run the next round with new testers: a fix often uncovers the next problem, and people who have seen the old version are no longer first-time users.

## Doing it from your coding agent

If recruiting and watching sessions is the part you will never get around to, [Hallway Test](https://hallwaytest.ai/) turns the whole loop into a tool call. Your coding agent describes the flow in one sentence, a fresh tester from our in-house team runs it on camera while thinking aloud, and the recording, timestamped transcript and friction points come back into the agent's context within 24 hours. It works in [Claude Code](https://hallwaytest.ai/connect-mcp), [Cursor](https://hallwaytest.ai/connect-cursor) and [Codex](https://hallwaytest.ai/connect-codex).

The first test is free in exchange for a 30-minute feedback call; the planned price after public launch is $29 per test. The guide to [usability testing for AI coding agents](https://hallwaytest.ai/learn/usability-testing-for-ai-coding-agents) explains where a real-user test fits next to automated checks, and [what is hallway testing](https://hallwaytest.ai/learn/hallway-testing) covers the method this one grew out of.

## What a usability test will not catch

A usability test shows where people get stuck. It does not check what they cannot see:

- **Security.** In Veracode's 2025 test of AI-generated code, [45% of samples failed security tests](https://www.veracode.com/blog/genai-code-security-report/) and introduced OWASP Top 10 vulnerabilities. A 2025 vulnerability in apps generated with Lovable came down to ["missing or insufficient Row Level Security (RLS) policies"](https://mattpalmer.io/posts/2025/05/CVE-2025-48757/). Review authentication and data access separately.
- **Load and data loss.** What happens with a thousand users at once or a failed payment webhook needs tests and monitoring.
- **Demand.** A flow can be easy to use and still solve a problem nobody has. That is a job for customer interviews.

## Sources

1. [Andrej Karpathy, post introducing vibe coding, X (February 2, 2025)](https://x.com/karpathy/status/1886192184808149383)
2. [Collins' Word of the Year 2025: AI meets authenticity as society shifts, Collins Dictionary (2025)](https://blog.collinsdictionary.com/language-lovers/collins-word-of-the-year-2025-ai-meets-authenticity-as-society-shifts/)
3. [Anna Kaley and Raluca Budiu, The Custodial Era of UX: Cleaning Up After AI, Nielsen Norman Group (2026)](https://www.nngroup.com/articles/ai-ux-debt/)
4. [AI, 2025 Stack Overflow Developer Survey](https://survey.stackoverflow.co/2025/ai)
5. [Marieke McCloskey, Turn User Goals into Task Scenarios for Usability Testing, Nielsen Norman Group (2014)](https://www.nngroup.com/articles/task-scenarios-usability-testing/)
6. [Amy Schade, Write Better Qualitative Usability Tasks, Nielsen Norman Group (2017)](https://www.nngroup.com/articles/better-usability-tasks/)
7. [Jakob Nielsen, Why You Only Need to Test with 5 Users, Nielsen Norman Group (2000)](https://www.nngroup.com/articles/why-you-only-need-to-test-with-5-users/)
8. [Steve Krug, Rocket Surgery Made Easy, table of contents (2010)](http://sensible.com/downloads/rsme-toc.pdf)
9. [The UX Research Incentives Report, User Interviews (2025)](https://www.userinterviews.com/blog/research-incentives-report)
10. [Jakob Nielsen, Thinking Aloud: The #1 Usability Tool, Nielsen Norman Group (2012)](https://www.nngroup.com/articles/thinking-aloud-the-1-usability-tool/)
11. [Jens Wessling, 2025 GenAI Code Security Report, Veracode (2025)](https://www.veracode.com/blog/genai-code-security-report/)
12. [Matt Palmer, CVE-2025-48757 (2025)](https://mattpalmer.io/posts/2025/05/CVE-2025-48757/)

## FAQ

**How many people do I need to test a vibe-coded app?**
Three to five per round. Jakob Nielsen's research found that five users uncover about 85% of the usability problems in a design, and that several small rounds with fixes in between find more than one big study. Steve Krug's lighter version is three people, one morning a month.

**Can I ask an AI agent to test my app instead?**
For obvious mistakes, yes: an agent clicking through the app finds dead links, broken steps and contradictory copy quickly. It does not get confused, hesitate or give up, and those moments are what a usability test exists to find. Use an agent while you build and real people before you ship a flow that matters.

**Do I need to pay testers?**
Strangers usually expect it. The User Interviews incentives report puts a typical remote consumer session at $80 an hour and recommends $100 an hour as a baseline for professionals. Friends and colleagues will test for free, but they are polite and tend to push through problems a customer would give up on.

**What should I test first in a vibe-coded app?**
The flow that decides whether the app succeeds: signing up and reaching the first useful screen, doing the core action for the first time, or paying. Test one flow at a time, with a clear finish line.

## Related

- [Home](https://hallwaytest.ai/index.md)
- [Pricing](https://hallwaytest.ai/pricing.md)
- [Book a free pilot](https://hallwaytest.ai/book.md)
