SolisReach
← Journal
Brand & UI/UX7 min read

Usability testing with five users and one afternoon

Written by the SolisReach team

Usability testing gets treated as a luxury reserved for teams with a research budget and a dedicated UX researcher. In practice, a lightweight version run in a single afternoon catches a large share of the problems a full formal study would find, and we recommend most small teams run it before, not instead of, a bigger investment later.

The barrier to running this kind of test is almost always perceived rather than real. Teams imagine they need a research lab, a recruiting agency, and a formal protocol before it's worth doing at all, and that perception keeps a genuinely cheap, fast diagnostic tool sitting unused. We push small teams specifically to run the lightweight version early and often, since the cost of running it is measured in hours, not weeks, and the findings routinely justify that cost many times over.

Five users is enough, and here's why

Usability research going back decades consistently shows that five users typically surface around 80 percent of the usability problems a much larger study would find, because most serious issues are obvious enough that they show up repeatedly with even a small sample. Beyond five, you mostly see the same problems again rather than new ones.

This isn't an argument against ever testing with more people, it's a specific claim about where the return on additional participants drops off sharply. The marginal sixth, seventh, and eighth user overwhelmingly confirm problems the first five already surfaced, rather than revealing meaningfully new ones. For a team deciding between testing with five people this week or waiting a month to recruit fifteen, the five-person test this week is almost always the better trade, since the problems it catches can start getting fixed immediately.

Recruit people who match your actual audience, loosely

They don't need to be perfect demographic matches. Five people roughly resembling your target user, colleagues from an adjacent team, friends of the founder who fit the general profile, are enough for this lightweight pass. Save precise recruiting for a later, more formal study if the product justifies the investment.

The perfectionism trap here is real: teams sometimes delay running any test at all because they can't find five perfectly matched participants on short notice, and the perfect test never happens because the ideal participants never materialize on schedule. A rough match run this afternoon beats a perfect match scheduled for three weeks from now, since the fundamental usability problems this test is designed to catch, confusing labels, unclear flows, buttons that don't look clickable, tend to affect almost anyone regardless of how closely they match the target demographic.

Give them a task, not a tour

Don't walk users through the product. Give them a specific task, "sign up and complete your first booking," and watch silently. The gap between what you assumed was obvious and what actually confused someone in real time is where the useful findings live, and narrating the interface yourself erases that gap entirely.

Write down what they do, not what they say

Users are notoriously unreliable narrators of their own confusion, often rationalizing a stumble after the fact. Note where they hesitated, clicked the wrong thing, or backtracked, regardless of what they say about it afterward. The behavior is the data; the commentary is context.

A user who fumbles for ten seconds looking for a button and then says "oh that was easy to find" once they finally locate it isn't being dishonest, they're doing what people naturally do: rationalizing a small struggle as soon as it resolves. We train ourselves to trust the ten seconds of hesitation over the after-the-fact comment, and to write down the specific moment of confusion in enough detail that it can be reviewed later without relying on memory.

Users are notoriously unreliable narrators of their own confusion, often rationalizing a stumble after the fact.

Set up the session so people forget they're being watched

We ask users to think aloud as they work, but keep our own presence minimal, no leading questions, no reassurance when someone struggles. It feels uncomfortable to sit quietly while someone visibly struggles with something you built, but jumping in to help erases the exact signal you're trying to capture.

The instinct to help is strong, especially when the person struggling is a colleague or a friend doing you a favor by participating. We've found it helps to state the rule out loud at the start of the session: "I'm not going to help you, and if you get stuck, that's useful information, not a failure on your part." Naming it upfront makes the silence feel like part of the process rather than an awkward withholding of help.

What to do with five sessions of notes

Look for patterns across at least three of the five sessions before treating something as a real problem rather than one person's quirk. A single afternoon run this way, done before a launch rather than after, has caught issues on several of our recent projects that would have otherwise shipped straight into a real user's first impression.

We also rank the patterns we find by severity before deciding what to fix, not just by frequency. A minor visual confusion that three of five users mentioned in passing matters less than a single user who couldn't complete the core task at all, even if that specific failure only showed up once. Frequency tells you how common a problem is; severity tells you how much it matters, and a good afternoon of testing needs both lenses applied to the same notes.

When to graduate to a bigger, more formal study

This lightweight version isn't meant to replace a proper research program forever. Once a product has meaningful usage and a dedicated research function becomes worth the investment, more rigorous methods, larger samples, more precise recruiting, statistical analysis of behavioral data, genuinely add value the afternoon version can't. The lightweight test earns its place specifically in the earlier stage, before that investment is justified yet.

The transition point isn't always obvious, and we've found it useful to frame it as a question of what kind of decision the research needs to support. Catching an obvious usability problem before launch is exactly what the afternoon version is for. Deciding between two subtly different onboarding flows with a marginal difference in conversion rate is a genuinely different kind of question, one that needs a larger sample and more statistical rigor than five sessions of qualitative observation can responsibly provide.

Start a project

Want this applied to your site?

We run a Core Web Vitals and SEO audit before quoting any performance marketing engagement, and we're happy to share what we'd find on yours.