# Real vs synthetic users for usability testing

> Synthetic users catch obvious UI mistakes fast. Real users get confused, hesitate and give up. When to use each, and how to combine both in one agent workflow.

Page: https://hallwaytest.ai/compare/real-users-vs-synthetic-users
Author: Kseniia Radova, Usability tester
Published: September 26, 2026. Last updated: September 26, 2026

**In short:** synthetic users, whether LLM personas or AI agents clicking through your site, are fast, cheap and good at catching obvious mistakes while you build. They do not get confused, hesitate, misread a label or give up. Real users do. Run synthetic checks during development; put one real person on any flow that has to convert before you ship it.

## What synthetic users are

"Synthetic user" covers two different things, and the difference matters.

**Interview simulators** prompt a language model to answer research questions as a persona: "You are a 34-year-old project manager evaluating invoicing tools." Tools of this kind are used for concept research and messaging rather than for clicking through screens.

**Agent testers** drive a real browser. An AI agent opens your site as a new visitor, tries to complete a task and reports what got in the way. Several startups and open-source MCP servers work this way. This is the kind that competes with a usability test, because the output looks like one: a list of friction points with screenshots.

Both are built on the same models, so both share the same strengths and the same blind spots.

## Where synthetic users help

Synthetic testing earns its place in four situations:

- **Before there is anything to click.** Naming, copy and concept questions can be pressure-tested against a model in minutes.
- **On every change.** An agent can re-run a flow after each commit. No person will do that.
- **For the obvious mistakes.** A dead link, a form that cannot be submitted, a step with contradictory instructions. Models are good at spotting what a careful reader would spot.
- **For desk research and hypotheses.** Nielsen Norman Group's evaluation of synthetic users concluded that they can support desk research and hypothesis generation in an unfamiliar domain, as long as nothing is decided on their answers alone.

The cost is close to zero and there is no scheduling, no consent form and no waiting. That is the whole appeal.

## Where they fail

In 2024, Maria Rosala and Kate Moran of Nielsen Norman Group [ran the same study with synthetic users and with real participants](https://www.nngroup.com/articles/synthetic-users/) and compared the answers. Their findings match what anyone who has watched real sessions will recognize:

- **Synthetic users are unrealistically optimistic.** Asked about an online course, a synthetic user reported finishing every module; real participants admitted stopping after three because of time and competing options.
- **They care about everything.** Real people rank their needs and ignore the rest. The synthetic answers listed every plausible concern with equal weight, so nothing in them told the team what to fix first.
- **They cannot explain why.** Real users tell stories: the situation they were in, what they tried, what they expected. Synthetic users produce generic lists. In the authors' words, "we are left knowing that these are common needs but have no idea how to address them."

Their conclusion is direct: "Real user research is essential. Synthetic users cannot replace the depth and empathy gained from studying and speaking with real people."

For usability testing specifically, the gap shows up in behaviour, not opinions. A model does not misread "Continue" as "Cancel", does not hover over a disabled-looking button for four seconds, does not open a second tab to check the price elsewhere, and does not close the tab out of frustration. Those moments are exactly what a usability test exists to find, and they are the moments that decide whether a flow converts.

There is a practical gap as well. A recording of a person getting stuck convinces a founder, a designer or a client in thirty seconds. A model's report of a "potential confusion" convinces nobody who did not already agree.

## Side by side

| | Synthetic users | Real users |
| --- | --- | --- |
| Gets confused, hesitates, gives up | No | Yes |
| Explains why in their own words | Generic reasons | Specific stories |
| Evidence you can replay | Text report, screenshots | Screen, webcam and voice recording with a transcript |
| Finds problems nobody anticipated | Rarely | Regularly |
| Turnaround | Seconds to minutes | Within 24 hours |
| Cost per flow | Near zero | Free first test with Hallway Test, then $29 planned |
| Matches your target audience | Only as well as the prompt | Only as well as the recruiting |
| Runs on every commit | Yes | No |
| Best moment to use | While building | Before shipping a flow that has to convert |

## A workflow that uses both

The two are not competing for the same slot, so the useful question is where each one goes. The guide to [usability testing for AI coding agents](https://hallwaytest.ai/learn/usability-testing-for-ai-coding-agents) places both next to automated tests.

1. **While building:** run an agent tester on every flow you touch. Fix what it reports the same way you fix a failing test.
2. **Before shipping a flow that matters:** sign-up, checkout, onboarding, anything with money or first impressions attached, put one real person on it and watch the recording.
3. **After the fix:** run the real test again with a different tester. Fresh eyes find what the fix uncovered.

Inside an MCP-based workflow this is two tool calls in sequence. The synthetic pass is whatever tool you already use. The real pass is Hallway Test's `create_usability_task`: your coding agent describes the flow in one sentence, a tester from our in-house team runs it on camera, and `wait_for_usability_result` hands back the recording, the timestamped transcript and the friction points into the same conversation. Setup takes one command in [Claude Code](https://hallwaytest.ai/connect-mcp), [Cursor](https://hallwaytest.ai/connect-cursor) or [Codex](https://hallwaytest.ai/connect-codex).

## Which one do you need right now?

Use a synthetic pass if the answer to all of these is yes:

- You are still changing the flow every day.
- You want to catch broken steps and contradictory copy, not judge the experience.
- Nobody outside the team will see this version.

Use a real tester if any of these is true:

- The flow is about to reach real customers.
- It involves money, sign-up or a first impression.
- You need to convince someone who has not seen the problem themselves.
- "It works for me and for the model" is the only evidence you have.

Most teams end up with both, in that order. What is worth avoiding is the third option: shipping a flow that only its author and a language model have ever completed.

## Sources

1. [Maria Rosala and Kate Moran, Synthetic Users: If, When, and How to Use AI-Generated Research, Nielsen Norman Group (2024)](https://www.nngroup.com/articles/synthetic-users/)
2. [Jakob Nielsen, Why You Only Need to Test with 5 Users, Nielsen Norman Group (2000)](https://www.nngroup.com/articles/why-you-only-need-to-test-with-5-users/)
3. [Joel Spolsky, The Joel Test: 12 Steps to Better Code (2000)](https://www.joelonsoftware.com/2000/08/09/the-joel-test-12-steps-to-better-code/)

## FAQ

**Are synthetic users accurate?**
For what a page says and whether a flow can be completed at all, often yes. For how a person actually behaves, no: Nielsen Norman Group's evaluation found synthetic users unrealistically optimistic, unable to prioritize, and unable to explain why they would make a choice. They describe an average user from their training data, not your users.

**Can an LLM replace usability testing?**
It can replace the part of usability testing that checks for obvious mistakes: broken links, contradictory copy, a step that cannot be completed. It cannot replace watching a person misread a button, hesitate, or abandon a flow, because a model does not get confused. Use it to reduce the number of real sessions you need, not to skip them.

**How many real users do you need?**
Three to five per round. Jakob Nielsen's research found that five users uncover about 85% of the usability problems in a design; after that, each extra session mostly repeats what you already know. Fix what the round found, then run the next round with new people.

**What does a real-user test cost with Hallway Test?**
The first test is free in exchange for a 30-minute feedback call. The planned price after public launch is $29 per test: one flow, one fresh in-house tester, up to 30 minutes, with the recording, transcript and friction points delivered into your coding agent.

## Related

- [Home](https://hallwaytest.ai/index.md)
- [Pricing](https://hallwaytest.ai/pricing.md)
- [Book a free pilot](https://hallwaytest.ai/book.md)
