Hallway testing is a quick, informal usability test: you take the next person who walks past, ask them to complete one task in your product without any help, and watch where they get stuck. It trades statistical rigor for speed, and it finds the obvious problems before your customers do.
Where the term comes from
The phrase was popularized by Joel Spolsky in 2000. Question twelve of his Joel Test, a checklist for software teams, asks “Do you do hallway usability testing?” and defines it in one sentence:
A hallway usability test is where you grab the next person that passes by in the hallway and force them to try to use the code you just wrote.
Spolsky claimed that doing this with five people uncovers 95% of the usability problems in your code. Jakob Nielsen had published the research behind that number a few months earlier: in Why You Only Need to Test with 5 Users he showed that the first five users find about 85% of the problems in a design, and that spending the same budget on three rounds of five users beats one round of fifteen, because you fix what the first round found before the next one starts.
The same method goes by other names. Corridor testing is the British variant of the term. Guerrilla testing is hallway testing taken outside the office, to a café or a co-working space, where the people you approach know nothing about your company. Wikipedia files all three under usability testing as the quick, cheap end of the spectrum.
What a hallway test looks like
A hallway test takes ten to fifteen minutes per person. The whole method fits in eight steps.
- Pick one task. Something a real customer would do, phrased as a goal, not as instructions: “buy a gift card for a friend”, not “click Shop, then Gift cards”.
- Find someone who has never seen the product. A colleague from another team, the person at the next desk, a friend. They do not need to match your target audience for a first pass.
- Give them the task and nothing else. No demo, no hints about where to click. If they ask a question, answer with “what would you do if I were not here?”
- Ask them to think aloud. What they expect to happen, what they are looking for, what surprised them. The words matter as much as the clicks.
- Do not help. Watching someone struggle with your work is uncomfortable. The struggle is the data.
- Note every hesitation, wrong click and workaround. Three or four friction points per session is normal. Write them down with the moment they happened.
- Fix the worst two. Usually a label, a button that looks disabled, a step nobody expected.
- Repeat with the next person. New rounds find the problems the fixes uncovered.
What it catches, and what it misses
| Problem | Does a hallway test find it? | Why |
|---|---|---|
| Confusing labels and wording | Yes | Anyone who reads the label hesitates |
| Buttons that do not look clickable | Yes | The tester looks for another way |
| Steps in the wrong order | Yes | The tester is surprised by what comes next |
| Missing feedback after an action | Yes | The tester clicks twice or asks “did it work?” |
| Broken or dead-end flows | Yes | The tester gives up |
| Domain-specific mistakes | Rarely | A random passer-by lacks the context |
| Long-term satisfaction and habit | No | One session shows first contact only |
| How common a problem is | No | Five people give a direction, not a percentage |
| Preferences of a specific audience | No | You tested whoever was nearby |
Hallway, guerrilla and formal usability testing compared
| Hallway test | Guerrilla test | Formal usability study | |
|---|---|---|---|
| Who tests | Whoever is nearby | Strangers in a public place | Recruited participants matching a profile |
| Where | Your desk or a call | Café, event, co-working space | Lab, moderated remote session or panel platform |
| People per round | 3 to 5 | 5 to 10 | 5 to 20 or more |
| Time to first result | Same day | A day | One to three weeks |
| Cost | Coffee | Coffee and vouchers | Recruiting fees, incentives, platform or agency |
| What you learn | Where people get stuck | Where people get stuck, with less bias | Task success rates, comparisons, audience-specific needs |
| Best moment | After every meaningful change | Before a launch | Before a redesign or a big bet |
How many people do you need?
Three to five per round, then fix, then another round. Nielsen’s reasoning still holds: the first tester finds most of the problems, the second finds mostly the same ones plus a few new, and after the fifth you are mainly watching people hit issues you already know about. More rounds with fresh people beat one large round, because each fix changes what the next tester will see.
When there is no hallway
Hallway testing assumes an office full of people who have not used your product. Remote teams, solo builders and small startups rarely have one, and software built with AI coding agents makes the gap wider: a feature can be designed, written and deployed in an afternoon without anyone outside the team touching it. The guide to testing a vibe-coded app with real users walks through a first test step by step.
The usual substitutes each cost something:
- Friends and family are available but polite. They rarely give up or say “this makes no sense”.
- Remote research platforms such as UserTesting recruit participants from a panel, which is right for a study and heavy for one flow. See the UserTesting alternative page for how the two approaches differ.
- Synthetic users, LLM personas or AI agents clicking through your site, are instant and free of scheduling, and they miss what makes a hallway test work: a person who can get confused. The real users vs synthetic users page goes through the trade-off.
Running a hallway test from your coding agent
Hallway Test turns the method into a tool call. Your coding agent describes the task in one sentence, for example “place an order as a guest and reach the confirmation page”, and a tester from our in-house team runs it on camera, thinking aloud. The screen recording, webcam reaction, timestamped transcript and the friction points come back into the agent’s context, ready to fix and to test again with the next person.
It works with any MCP-compatible client. In Claude Code it is one command:
claude mcp add --transport http hallwaytest https://mcp.hallwaytest.ai/mcp
Setup guides: Claude, Cursor and Codex. The first test is free in exchange for a 30-minute feedback call; the planned price after launch is $29 per test. Details on the pricing page.