In short: a coding agent can write a flow, run the tests and even click through the result in a real browser. It cannot tell you where a person who has never seen the flow gets stuck, because it wrote the flow. Usability testing for AI coding agents adds one step to the agent loop: a real person tries the flow, and the recording and transcript come back to the agent as evidence to fix.
Why the agent loop needs a usability step
AI coding tools are now the default way software gets written. In the 2025 Stack Overflow Developer Survey, 84% of respondents said they use or plan to use AI tools in development, and 51% of professional developers use them daily. Google’s 2025 DORA research puts AI use at work at 90%.
Trust has not kept pace. More developers distrust the accuracy of AI tools (46%) than trust it (33%), and the most common frustration, named by 66%, is “AI solutions that are almost right, but not quite.” DORA’s summary is shorter: “AI doesn’t fix a team; it amplifies what’s already there.”
“Almost right” is the usual state of a screen an agent has just built. The code compiles, the tests pass, the page renders. Whether a customer understands it is a separate question, and nobody in the loop is in a position to answer it:
- The agent is the ultimate insider. It wrote the spec, the copy and the code. Nielsen Norman Group’s warning about teams testing their own products applies in full: “once you know how something works, you cannot reliably simulate not knowing.”
- The developer is the second insider. They wrote the prompt and know what the screen is supposed to do.
- The loop is fast. A flow can go from prompt to production in an afternoon without anyone outside the team touching it. Hallway testing, which used to catch the obvious problems, needs an office full of people who have not seen the product, and most teams building with agents do not have one.
What an agent can already check by itself
A lot, and it should keep doing it.
- Tests answer whether the code does what the spec says. Agents write and run them as part of the task.
- Browser automation over MCP answers whether the flow works in a real browser. Playwright MCP lets a model interact with web pages “through structured accessibility snapshots”, and GitHub enables it by default for its Copilot coding agent. Google released Chrome DevTools MCP in September 2025 because coding agents “are not able to see what the code they generate actually does when it runs in the browser. They’re effectively programming with a blindfold on.”
- Synthetic users, AI agents that play a first-time visitor, answer whether a model can complete the flow, and they spot obvious mistakes such as a dead link or contradictory copy. The real vs synthetic users comparison covers what they catch and what they miss.
All three run in seconds or minutes and can run on every change. None of them involves a person.
What it cannot check
Quality assurance and usability testing ask different questions. NN/g’s example is an expense reimbursement: QA confirms the money arrives, and “whether the user enjoys that reimbursement process is irrelevant to quality assurance.” The same split applies to anything an agent verifies about its own work.
An agent driving the browser already knows that the label means “continue”, where the button is and what the next step will ask for. A first-time user does not, so they:
- misread a label and go back a step;
- miss a button below the fold on a laptop screen;
- hesitate when the price appears later than they expected;
- abandon a form that asks for a phone number without saying why.
Models acting as users do not close the gap. NN/g found that synthetic users “often provide shallow or overly favorable feedback”, and its guidance on AI in research is explicit: “Avoid using AI tools to moderate usability tests; AI research tools are not (yet) capable of actually knowing what users are doing.”
Three layers of testing for agent-built software
| Layer | Question it answers | How | When | Time |
|---|---|---|---|---|
| Tests | Does the code do what the spec says? | Unit and end-to-end tests the agent writes | Every change | Seconds |
| Browser checks | Can the flow be completed in a real browser? | Playwright MCP, Chrome DevTools MCP, agent testers | Every change to the UI | Minutes |
| Real users | Where does a person who has never seen it get stuck? | A human tester, recorded while thinking aloud | Before shipping a flow that matters, and after fixing it | Within 24 hours with Hallway Test |
The first two layers keep the build honest. The third is the only one that measures the product against someone who did not build it.
How to put a real user in the loop
The goal is a real-user test that behaves like any other tool call: the agent starts it, waits for it and acts on the result.
- Pick the flow and the finish line. One flow per test, with a clear end state: “sign up with email and reach the dashboard”, “place an order as a guest and reach the confirmation page”.
- Put it where a stranger can open it. A staging or preview URL rather than localhost, plus any test credentials the tester needs.
- Dispatch the test from the agent. With Hallway Test, the agent calls
create_usability_taskwith the flow in one sentence, and a fresh tester from our in-house team runs it on camera, thinking aloud. If you are not sure which flows to test first,plan_usability_test_suitescans the project and drafts 5 to 10 critical flows ranked by risk. - Let the agent read the evidence.
wait_for_usability_resulthands back the recording as soon as it is ready. The agent then reads the click and navigation log, the timestamped think-aloud transcript, screenshots from key moments and the console and network logs, and lists each friction point with the moment it happened. - Fix, then test with someone new. The agent proposes changes, you review them, and the next test goes to a different tester. Jakob Nielsen’s research found that five users uncover about 85% of usability problems, and that small rounds with fixes in between beat one large study.
Setting it up in Claude Code, Cursor and Codex
Hallway Test is a remote server for the Model Context Protocol, the open standard that Claude Code, Cursor and Codex use to connect to external tools. Setup is one command or one configuration entry, plus a Google sign-in:
- Claude Code:
claude mcp add --transport http hallwaytest https://mcp.hallwaytest.ai/mcp(guide) - Cursor: one entry in
mcp.json(guide) - Codex:
codex mcp add hallwaytest --url https://mcp.hallwaytest.ai/mcp(guide)
The first test is free in exchange for a 30-minute feedback call; the planned price after public launch is $29 per test. The pricing page lists what every test includes.
How this differs from UserTesting’s MCP server
In September 2026, UserTesting released an MCP server that lets teams “recruit participants, create studies, and launch tests from Claude, ChatGPT, Figma Make, and other AI clients”. It is in early access and brings a full research platform into AI chat and design tools, which suits research and design teams running studies with a specific audience.
Hallway Test is built for the coding agent’s loop instead: one flow, one fresh tester, and the evidence delivered into the same session where the code gets fixed. The UserTesting alternative page compares the two in detail.
Sources
- AI, 2025 Stack Overflow Developer Survey
- Nathen Harvey and Derek DeBellis, Announcing the 2025 DORA Report: State of AI-Assisted Software Development, Google Cloud (2025)
- Therese Fessenden, Dogfooding vs. QA vs. User Research, Nielsen Norman Group (2026)
- Maria Rosala and Kate Moran, Synthetic Users: If, When, and How to Use AI-Generated Research, Nielsen Norman Group (2024)
- Kate Moran and Maria Rosala, Accelerating Research with AI, Nielsen Norman Group (2024)
- Jakob Nielsen, Why You Only Need to Test with 5 Users, Nielsen Norman Group (2000)
- Mathias Bynens and Michael Hablich, Chrome DevTools (MCP) for your AI agent, Chrome for Developers (2025)
- Playwright MCP, Microsoft on GitHub
- What is the Model Context Protocol (MCP)?, modelcontextprotocol.io
- UserTesting MCP Server, product page