SolisReach
← Journal
Brand & UI/UX5 min read

Why we test navigation with five users before we ship it

Written by the SolisReach team

Usability testing research going back decades consistently finds that testing with five users uncovers the large majority of a navigation structure's real problems, and each additional user beyond that finds diminishing new issues. We've validated this ourselves across dozens of client projects: five focused sessions almost always surface the problems worth fixing before launch.

The research behind this number, originally from usability researcher Jakob Nielsen, is often misquoted as meaning five is always enough for any kind of testing. It isn't a universal rule, it's specific to qualitative usability problems in a relatively homogeneous user group, which is exactly the situation most navigation testing sits in, and part of why we still explain the reasoning to clients rather than just citing the number as an unquestioned rule of thumb.

What we're actually testing for

Not "do people like this," which is a weak, unreliable signal. We give each participant a specific, realistic task, find the pricing page, locate the return policy, start the signup flow, and watch where they hesitate, backtrack, or give up. The navigation structure itself, not the visual design, is what we're evaluating, so we deliberately choose tasks that depend on the person finding their way rather than judging aesthetics.

We deliberately avoid leading the participant during the task, resisting the urge to clarify or hint even when it's obvious they're headed the wrong way, since the moment we intervene, we lose the exact data point we're there to collect: whether the navigation on its own, with no help, gets someone where they need to go.

How we actually recruit the five

We aim for people who roughly match the target audience but have never seen the site before, since anyone internally familiar with the navigation is the wrong tester; they already know where things are. Five strangers who match the target demographic, even recruited informally, produce more useful signal than a larger group of colleagues who already know the answer to every task we'd give them.

For a genuinely niche B2B product, we sometimes have to relax the audience-match requirement slightly, since finding five genuine target users willing to sit for a test on short notice isn't always realistic. In that case, we recruit people with a reasonably similar level of general web fluency and note the limitation explicitly in how we report the findings, rather than presenting the results with more confidence than the recruiting process actually supports.

The pattern that shows up almost every time

Across most of these tests, at least two of the five participants get stuck on the same specific point in the navigation, independently, without prompting each other. That convergence is the signal: a problem two unrelated people both hit independently is a real structural issue, not a fluke, and it's almost always something the internal team, too close to their own navigation, never would have flagged on their own.

We record every session, with consent, specifically so we can show the client the actual moment of hesitation rather than just describing it secondhand. Watching a real person pause and backtrack on video is far more persuasive to a stakeholder than a bullet point in a findings report, and it tends to end debates about whether a fix is really necessary.

What we do with the findings

We fix anything two or more of the five participants struggled with before launch, treat single-person struggles as worth watching but not necessarily worth a pre-launch fix, and re-test the specific fixed flow with two or three new participants if the change was significant enough to risk introducing a new problem.

This threshold, two or more out of five, keeps the process fast and decisive rather than turning every session into an open-ended debate about which findings matter. A clear, pre-agreed rule for what counts as worth fixing removes a lot of the back-and-forth that otherwise slows a findings review down.

We fix anything two or more of the five participants struggled with before launch, treat single-person struggles as worth watching but not necessarily worth a pre-launch fix, and re-test the specific fixed flow with two or three new participants if the change was significant enough to risk introducing a new problem.

Why this is worth the two or three days it costs

A navigation problem caught in a two-day pre-launch test costs a couple of days to fix. The same problem discovered after launch, once it's already cost real conversions and real user frustration, costs far more to diagnose (since analytics rarely explains why someone left) and to fix under pressure. Five short user sessions are cheap insurance against a much more expensive discovery process later.

We've had more than one client initially skeptical of spending two or three days on this, only to become the strongest advocate for keeping it in every future project once they've watched a single session where a real person got visibly, unmistakably stuck on something the internal team had never once questioned.

Remote testing versus in-person, and when each is worth it

Remote sessions over a video call, with the participant sharing their screen, work well for most navigation testing and are considerably easier to schedule than in-person sessions, since they don't require anyone to travel. We default to remote for most projects for exactly this reason, and it hasn't meaningfully weakened the quality of what we learn.

In-person testing still has a real edge for physical products or in-store digital experiences, a kiosk, a point-of-sale screen, where body language and physical hesitation carry information a screen share can't fully capture. We reserve in-person sessions for these specific cases rather than defaulting to them for a standard website navigation test where the extra cost and scheduling friction wouldn't buy much additional insight.

One practical note on remote sessions: we always ask participants to use their own device rather than a loaner or a shared testing rig, since navigating with a genuinely familiar keyboard, mouse, or touchscreen produces more natural behavior than navigating on unfamiliar hardware, which can itself introduce hesitation that has nothing to do with the actual navigation being tested.

We also always run at least one of the five sessions on mobile specifically, even for a product primarily used on desktop, since navigation problems on a smaller screen are frequently different in kind, not just degree, from the ones that show up on a larger one, and skipping mobile entirely misses a meaningful share of real usage for most consumer-facing sites.

Start a project

Want this applied to your site?

We run a Core Web Vitals and SEO audit before quoting any performance marketing engagement, and we're happy to share what we'd find on yours.