Design hiring guide

Can you really test a prototype with five users?

Five people will find most of the serious usability problems in a single flow, and five people will tell you nothing reliable about preference, conversion or how common anything is. The number is quoted constantly and misapplied almost as often, because it answers a discovery question rather than a measurement one. Knowing which of those you are asking decides whether five is generous or absurd.

What can five people actually tell you?

That something is broken, and roughly why. If three of five people cannot find the action, you do not need a larger sample to act, because the problem is already demonstrated and a bigger study would only tell you how many others share it. Discovery of defects saturates quickly, which is the whole basis of the small-sample approach.

What five people cannot tell you is anything with a number attached. They cannot tell you which of two designs converts better, what proportion of users would pay, or whether a change improved anything, because the variation between five individuals swamps any effect you care about. Those are experiments, and they need traffic rather than sessions.

The practical rule is that small-sample testing answers what and why, and instrumentation answers how many. Teams get into trouble by taking a preference expressed by three of five participants into a board meeting as though it were a finding.

When are five people not enough?

When your users are not one group. Five is a per-audience number, not a total. If your product serves both administrators and occasional end users, or both clinicians and patients, you need five of each, because their mental models differ and the problems one group hits are invisible to the other.

It is also not enough when the task is rare or specialised. Testing an annual reconciliation process with five people who have never done a reconciliation produces confident nonsense. Sample size stops being the constraint and recruitment quality becomes everything.

And it is not enough when the flow is long. Five people through a fifteen-step onboarding journey will surface problems in the first three steps and get progressively less useful, because early failures change how they approach everything after. Test long journeys in segments with different participants rather than dragging the same five through the whole thing.

What you want to knowRight methodPeople needed
Where does this flow breakModerated task sessionFive per distinct audience
Do people understand this conceptComprehension test on a rough artefactFive to eight
Which of two designs performs betterLive experiment with instrumentationEnough traffic, not enough people
Do people find this feature at allFirst-click or findability test, unmoderatedTwenty to fifty
Is the wording clear to non-native speakersModerated session with that audience specificallyFive of that audience
How common is this problemAnalytics or support ticket analysisYour whole user base

How do you run a session that produces something useful?

Give a task, not a tour. The single biggest difference between a useful session and a demonstration is whether the participant is trying to achieve something or being shown something. Say what they are trying to do in their own terms, hand over control, and stop talking.

Write the tasks so they contain no interface vocabulary. Asking someone to open settings and enable notifications tells you nothing, because you have already given them the route. Ask them to make sure they will hear about it when a delivery arrives, and watch where they go. If you cannot phrase the task without naming your own buttons, the task is testing your memory rather than their understanding.

Then say almost nothing for the rest of the session. When asked what to do next, return the question: what would you do if you were on your own. Silence is uncomfortable and it is where the findings are. The most valuable minute in most sessions is the one where a participant is stuck and the facilitator has not rescued them.

Which questions ruin the results?

Anything asking someone to predict their own behaviour. Would you use this, would you pay for this, would you recommend this. People are consistently generous and consistently wrong about future behaviour, especially to a person who has just spent half an hour being nice to them.

Anything containing the answer. Was that easy to find, do you like the new layout, is this clearer than before. Leading questions get agreement in a research setting almost every time, and the resulting quote gets pasted into a slide as evidence.

Preference questions between designs also mislead, because people rationalise. If you must ask which they prefer, ask after they have completed the tasks in both, and treat the answer as a tie-breaker rather than a result. What they did carries information. What they say they prefer mostly reflects what they saw last.

Who should be in the room from your side?

One facilitator, and as many silent observers as you can get. The observers are the point. A finding an engineer watched happen needs no persuasion, and a finding delivered as a written recommendation gets debated for a fortnight. Getting the people who will build the thing to watch two sessions each is worth more than any report.

Give observers one job: write down what they saw, not what they concluded. The discipline is to record that the participant scrolled past the panel three times, rather than that the panel needs to be bigger. Solutions belong in the debrief, and observers who arrive with solutions stop noticing.

Debrief within the hour, before memory rewrites itself. Ten minutes, everyone lists what surprised them, and the list gets ordered by how many participants hit it. That ordered list is the deliverable, and it is more useful than a thirty-page report nobody opens.

What can you run this afternoon?

Take your current prototype, write three tasks in the user's own language, and stop five colleagues who have never seen the project. Colleagues are a compromised sample and they will still find the worst problems, because those problems are structural rather than subtle. Twenty minutes each, one facilitator, one observer.

The bar to clear is whether anything surprised you. If nothing did, you either tested with people who already knew the answer, or gave tasks that named your own buttons, or rescued participants too early. All three are facilitation problems rather than evidence that the design is fine.

Then repeat the same tasks with genuine users before deciding anything expensive. The colleague test tells you whether the prototype and the tasks work. The real test tells you whether the design does, and those are different afternoons.

Common questions

Is testing with five users really enough?
Enough to find most serious usability problems in a single flow, and not enough to measure anything. If three of five participants cannot find an action, the problem is demonstrated and a larger sample would only tell you how many others share it. Comparing two designs, estimating conversion or judging whether a change helped all require traffic and instrumentation instead.
When do you need more than five participants?
When your users form distinct groups, because five is a per-audience number and administrators, occasional users, clinicians and patients each hit different problems. Also when the task is rare or specialised, where recruitment quality matters more than count, and when the journey is long, since early failures distort everything a participant does afterwards. Long journeys are better tested in segments with different people.
How do you write a usability test task?
State the goal in the participant's own words and include no interface vocabulary. Asking someone to open settings and switch on notifications gives away the route and tests nothing. Asking them to make sure they will hear about it when a delivery arrives leaves the route to them, which is the part under test. If a task cannot be phrased without naming your buttons, it is testing recall rather than understanding.
What questions should you avoid in user testing?
Anything asking someone to predict their behaviour, such as whether they would use or pay for something, because people are generous and unreliable about their own future actions. Anything containing its own answer, such as asking whether something was easy to find. Preference questions between designs also mislead, since answers tend to reflect what was seen most recently rather than what worked better.
Who should watch a usability session?
One facilitator and as many silent observers as possible, especially the engineers who will build the thing, because a problem someone watched happen needs no persuading. Give observers the single job of recording what they saw rather than what they concluded, and debrief within the hour while memory is intact. An ordered list of what surprised people beats a long written report.

More on UI/UX designers

Let’s create something out of this world together.

Have a project in mind? Contact us for expert design and development solutions. Let’s discuss how we can help grow your business.

Azaadi Offer

Claim a free security assessment

Until 31 August we're covering the cost of a full vulnerability assessment and penetration test. Mention it in your message and we'll scope it with you.

  • Web application testing, authenticated and unauthenticated
  • Mobile application testing across iOS and Android
  • External network and infrastructure assessment
  • Manual exploitation by engineers, not scanner output

Testing and the report are free. Fixing what we find is quoted separately, with no obligation to accept.

Read the full offer

Tell us what you are trying to build and we will tell you plainly whether we are the right people for it. Book a call with an expert to work through the detail, or ask for a fixed quote if the scope is already clear. No obligation either way.

Four fields is all we need to get started.

Fastnexa Logo

© 2026 fastnexa. All rights reserved.