Sevginur Ak Parlak

Sep 3, 2026

  • 13 min read

Usability Testing With 5 Users: How I Run It on a Prototype

Most teams tell me they have already tested their prototype. What they usually mean is that they showed it to a dozen people at a meetup, everyone said it looked great, and 2 people asked when it launches. Take a typical case: 5 proper sessions with people who match the customer, and most of them cannot complete the main task without help. Nothing had been tested. The prototype had been demonstrated.

Usability testing with 5 users is the method that finds this before the developers do. The details decide whether you learn something or confirm what you already believed: who you recruit, how you phrase the task, whether you stay quiet, and what you do with the notes.

This post is the practical version of how I run it on a clickable prototype, for founders and product teams who want to do it themselves or want to know what they are buying. It is the method inside the structure validation work we do for websites and mobile apps.

Why 5 users is enough, and when it is not

5 users per role find the majority of usability problems in a flow. This is not my number, it comes from research repeated many times since the 1990s. The 1st user finds about 30% of the problems. Each new user finds fewer new ones. By the 5th, you are watching the same 8 problems again, and the 6th session teaches you very little.

There are 2 conditions. The 5 users must match the people who will use the product, and they must be doing 1 role. If your product has a buyer and a seller, that is 5 of each. If it has an admin and a member, 5 of each.

What 5 users cannot do is give you a percentage. If 3 of 5 fail a task, I do not write "60% of users fail". I write that the task has a problem worth fixing, which is the only decision the test needs to support.

What to test: fidelity and the 3 to 5 tasks

I test on a clickable prototype in Claude, sometimes greyscale, with real labels and working navigation for the flows under test. Not a static wireframe, because users cannot get lost in a picture.

Before I recruit anyone I write down 3 to 5 tasks. Each task is 1 thing a real user would come to the product to do, and it has a clear end point I can observe. "Create a project and invite 1 colleague" has an end point. "Explore the dashboard" does not. If I cannot say in 1 sentence what success looks like, the task is not ready.

I also decide what I am measuring before the sessions: completed without help, completed with help, or gave up, plus every hesitation longer than 3 seconds.

Recruiting the 5 users

The most common mistake in usability testing with 5 users is the recruiting, not the moderation. 5 sessions with the wrong people produce confident, wrong findings.

I write a 5 question screener. Who they are, what they do, what tool they use for this job today, how often they do the task, and 1 question that disqualifies people who work in design or product. Then I look for people in 3 places: the client's own waiting list or customer base, communities where the target user already talks, and a research panel when the profile is narrow.

Friends, colleagues and investors are excluded. They know too much or they want the product to succeed. Nobody who has seen the prototype before, because the 2nd time through a flow tells you nothing about the 1st.

I book 6 for 5. Somebody always cancels. Sessions are 45 minutes, on a video call with screen sharing, or in person for a mobile app where I want to see the hands. I spread them over 2 days.

Writing tasks that do not lead

The words in the task decide what you learn. If the task contains the label on the button, the participant is not finding anything, they are matching words.

Each task starts with a situation, not an instruction. It uses the participant's vocabulary, ends with an action I can observe, and does not contain any word from the navigation.

I write the tasks on 1 page, in the order a real user would meet them, with the success condition next to each. That page is the script, and I read it out loud exactly as written, every session, so the 5 sessions are comparable.

Moderating without leading

The hardest part of moderation is silence. A participant stops, looks at the screen, and every instinct says to help. I count to 10 before saying anything. Most of the time they find it, and where they looked first is the finding.

I ask 3 things and almost nothing else. "What are you looking for?" when they pause. "What did you expect to happen?" after they tap something. "What would you do next?" when they stop. Never "would you click here", never "do you like this", never an explanation of what the screen is for.

At the start I say 1 sentence that changes the whole session: "We are testing the design, not you. If something is confusing, that is a problem with the design, and it is exactly what we want to find." Then I ask them to think out loud.

When a participant asks a question, I answer with a question. "Can I click this?" becomes "What do you think would happen?" If they are stuck for more than 2 minutes, I help, and I write down that I helped.

Note taking and the 2 person setup

I do not moderate and take notes at the same time. 1 person moderates, 1 person watches and writes. If you are alone, record the session and take notes from the recording the same day.

The note taker uses 1 sheet per participant with 1 row per task. For each task: completed, completed with help, or gave up. The time it took. Every hesitation, with a timestamp and what the participant said. Every wrong click, with the label they clicked instead. Opinions about the visual design go in a separate column that I read last.

After each session, the 2 of us spend 10 minutes writing the 3 biggest things we saw. After 5 sessions, those 15 lines are the first draft of the findings.

Turning findings into structure changes

I put every observation on a card and group them by cause, not by screen. 3 participants hesitating on 3 different screens because the word "Workspace" means nothing to them is 1 finding, not 3. Then I rate each finding: a blocker stops the core task, a major finding causes a wrong click or a request for help, a minor one is a hesitation.

Then for each finding I write what changes in the prototype. A label, an order of steps, a section moving in the navigation, a missing empty state. I make the change in the wireframe, and if it is big I test it again with 3 more people. In a typical validation, expect the structure of 20 to 40% of the screens to change after the 5 sessions.

The report the team gets has 1 page per finding: the evidence as a 20 second clip, the severity, the change, and the screen before and after. Founders read 1 page. They act on the clip.

Where AI tools help and where they do not

AI tools are useful before and after the sessions. I use them to draft the screener and the task script, to transcribe 5 recordings, to pull every hesitation with its timestamp out of the transcripts, and to group observations by cause as a first pass.

What they cannot do is sit in the session. AI does not see the participant's hand move to the wrong tab and stop. It cannot decide when to stay silent and when to help. And synthetic users, where an AI pretends to be your customer, produce an average of every user of every product, which is not your user.

Mistakes I see in app navigation design and testing

Testing with the team. Colleagues know where everything is. 5 internal sessions produce 0 structural findings and a false sense of confidence.

Demonstrating instead of testing. The founder shares the screen, walks through the prototype, and asks what people think. Everyone says it looks great. Nothing was tested.

Tasks that contain the answer. "Click Projects and create a project" tests whether the participant can read.

Helping at the 1st pause. The moderator jumps in after 3 seconds and never learns where the participant would have looked.

Asking what they would do instead of watching what they do. People predict their own behaviour badly. Only what they do is useful.

Fixing during the sessions. Changing the prototype between session 2 and 3 means the 5 sessions are no longer comparable.

FAQ

Is usability testing with 5 users enough?
For 1 user role and 1 set of flows, yes. 5 people who match the target user find the majority of the usability problems, and the 6th session mostly repeats what the first 5 showed. If the product has 2 roles that matter, test 5 of each. If you need a percentage rather than a list of problems, you need a different method with far more people.

How do I run a usability test on a prototype?
Write 3 to 5 tasks with a clear success condition, make the prototype clickable for those tasks, recruit 5 people who match your customer and have not seen it, and run 45 minute sessions where you read the task and stay silent. Have a 2nd person take notes. Group the notes by cause, rate them, change the prototype.

How long should a usability test session be?
45 minutes. 5 minutes of introduction and warm up, 30 minutes for 4 to 6 tasks, 10 minutes of follow up questions. Longer than 60 minutes and participants get tired and start guessing.

How do I find participants for usability testing?
Your own waiting list or customer base first, then communities where the target user already talks, then a research panel when the profile is narrow. Use a 5 question screener and exclude friends, colleagues, investors and anyone who works in design or product. Book 6 to end up with 5.

What do I do with the results of a usability test?
Group every observation by cause, not by screen, and rate each finding as blocker, major or minor. Write the change to the structure for each one, make it in the prototype, and test the big changes again with 3 people. The output is a changed prototype, not a document.

Working with us

We run usability testing with 5 users as part of structure validation for websites, mobile apps and SaaS products, in 2 weeks from planning to the changed prototype. We write the tasks, recruit, moderate, take notes and turn the findings into a revised clickable prototype in Figma, with a 20 second clip for every change. We work in Figma, Claude, Slack and Loom, async across timezones.

If you have a prototype and you are not sure whether people can use it, that is exactly what 5 sessions are for. Book a free 15 minute intro call and show us the flow. We will tell you what we would test first, even if you run the sessions yourself.

Get a senior design partner on your team by next week.

Book a free 15-minute intro call. We'll review your product live and tell you exactly what we'd fix first, yours to keep either way.

Get a senior design partner on your team by next week.

Book a free 15-minute intro call. We'll review your product live and tell you exactly what we'd fix first, yours to keep either way.

Get a senior design partner on your team by next week.

Book a free 15-minute intro call. We'll review your product live and tell you exactly what we'd fix first, yours to keep either way.