Builders Camp

Other Guides

How to Run a Usability Test (When You're the PM, Not the Researcher)

To run a usability test as a PM, name the decision it feeds, recruit about 5 people per user group, give them realistic tasks and ask them to think aloud, then rate each problem by frequency, impact and persistence and turn the top ones into roadmap items. Own the decision, the questions and the follow-through; delegate the screener, moderation and detailed analysis when a researcher is available.

What is a usability test, and why should a PM run one?

A usability test puts a real or prototype product in front of representative users, gives them tasks, and watches where they succeed, hesitate and fail. The standard definition comes from ISO 9241-11:2018, quoted in the NIST glossary: usability is "the extent to which a product can be used by specified users to achieve specified goals with effectiveness, efficiency, and satisfaction in a specified context of use." A test measures those three things for one group of users on one set of tasks.

It is also cheaper than most PMs expect. Jakob Nielsen's analysis in Why You Only Need to Test with 5 Users found that a single test user reveals almost a third of the usability problems in a design, and a first study with 5 users finds about 85%. His recommendation: "The best results come from testing no more than 5 users and running as many small tests as you can afford."

The six stages below are plan, recruit, write the guide, run the sessions, analyse, and report. For each one, the question for a PM is what you must own and what you can hand to a researcher or designer.

Which parts of a usability test should the PM own?

Stage PM owns Can be delegated
Plan The decision the test feeds, the research questions, which flows to test Detailed test protocol, consent forms
Recruit Who counts as a representative user Screener, scheduling, incentives
Session guide Task scenarios match real user goals Wording, warm-up questions, pilot run
Sessions Watching at least two sessions live Moderation, note-taking, recording
Analysis Agreeing severity for the top issues Coding notes, affinity mapping, counts
Report Turning findings into decisions and roadmap items Highlight clips, detailed issue log

If there is no researcher, you own all of it, and the lightweight version of each stage below still works.

How do you plan a usability test?

Start from the decision, not the feature. "Test the new prescription refill flow" is a topic. "Decide whether the refill flow can ship next sprint or needs another design round" is a decision, and it tells you what to test and what a good result looks like.

Take an original example: a pharmacy app is redesigning how patients request a prescription refill. A plan for that test fits on one page:

  • Decision: ship the new refill flow next sprint, or revise it first.
  • Research questions: Can patients find the refill option from the home screen? Do they understand which prescriptions are eligible? Do they know when the refill will be ready?
  • Tasks: request a refill of a regular medication; check when it will be ready; change the pickup pharmacy.
  • Success criteria: at least 4 of 5 participants complete the first task without help; nobody leaves unsure whether the request went through.
  • Participants: 5 patients who refill at least one prescription a month, including at least two over 65, because older patients refill most often in this hypothetical app.
  • Format: moderated, remote, 45 minutes, clickable prototype.

The lightweight test plan challenge walks through the same fields on a different scenario if you want to practise before running one for real.

How do you recruit participants for a usability test?

Recruit for behaviour, not demographics. A screener question like "How often did you refill a prescription in the last three months?" filters better than an age bracket, because it describes the task you are testing.

Five per round is the default, with one exception Nielsen names directly: "You need to test additional users when a website has several highly distinct groups of users." In the pharmacy example, patients and pharmacy staff use different sides of the product and need separate groups. Avoid recruiting colleagues, friends or power users from your beta list; they know too much and will complete tasks that real users fail. For participants who are less confident with technology, run a short setup call before the session so the first ten minutes are not spent on screen sharing.

How do you write tasks and a session guide?

A session guide has four parts: an introduction that explains you are testing the product, not the person; a few warm-up questions about how they handle the task today; the task scenarios, read aloud one at a time; and wrap-up questions.

The tasks do most of the work. Write them as realistic scenarios with a goal, not as instructions that name the interface:

Weak task Stronger task
"Click Refill and select your medication." "You have three days of your blood pressure tablets left. Get more before they run out."
"Do you like the new status page?" "You requested a refill this morning. Find out whether you can pick it up after work today."
"Try the pharmacy switcher." "You are staying with family next week. Arrange to collect your refill from a pharmacy near them."

Pilot the guide with one colleague first. Most broken tasks (a scenario that gives away the answer, a prototype link that does not work) show up in that one run.

How do you moderate a usability session?

Ask participants to think aloud, then mostly stay quiet. Nielsen called it the method he values most in Thinking Aloud: The #1 Usability Tool: "Thinking aloud may be the single most valuable usability engineering method." His summary of how to run one is three steps: recruit representative users, give them representative tasks, and let the users do the talking.

A few habits keep the data clean:

  • Use neutral prompts. "What are you looking for?" and "What did you expect to happen?" instead of "Did you see the button at the top?"
  • Let people fail. Only step in when a participant is stuck long enough that the rest of the session is at risk, and note exactly where it happened.
  • Answer questions with questions. "What do you think it does?" tells you more than explaining it.
  • Reset between sessions. Clear test accounts, carts and history so the next participant starts from the same place.

Invite engineers and designers to observe. One engineer watching a patient give up on the refill flow does more for the fix's priority than a slide deck will.

How do you analyse findings and rate severity?

Capture every observation in the same format (participant, task, what happened, quote), then group observations into issues. For each issue, rate severity. Nielsen Norman Group's severity ratings combine three factors: frequency (is it common or rare?), impact (is it easy or hard to overcome?), and persistence (does it keep bothering users once they know about it?). NN/g then rolls them into a single 0 to 4 score, from 0 (not a usability problem) through 1 (cosmetic) and 2 (minor) to 3 (major, high priority) and 4 (usability catastrophe, fix before release).

Frequency is the only factor you can count; impact and persistence need judgement, so agree them with the designer rather than scoring alone. The AI for usability test analysis guide covers how to use an AI assistant for the counting and the quote retrieval without letting it decide severity.

How do you turn usability findings into roadmap items?

A usability report that lists 30 issues gets read once. A report that leads with the three most severe issues, each tied to a decision, gets acted on. For each top issue, write the finding, the evidence (a count, a short quote, a clip), the severity, and the proposed action:

Finding Evidence Severity Action
Patients cannot tell whether the refill request was sent 3 of 5 went back to check; two phoned the pharmacy 4 Fix before release: add a confirmation screen and a status message
"Eligible" label is unclear 2 of 5 asked what it meant 2 Rewrite copy in the current sprint
Pharmacy switcher is hard to find 1 of 5 could not find it 3 for travellers, low frequency overall Backlog item; add to next round's tasks

Speak the business's language to executives: tie each severe issue to the metric it threatens, such as support calls, abandoned refills or app ratings. Include what worked too, so the team does not redesign the parts users handled well. Then close the loop: retest after the fixes with a new round of 5, because Nielsen's point is that a second round checks whether the fixes worked and finds the problems the first round's issues were hiding. Keep a running log of issues and fixes so the next test starts from what you already know.

When is a usability test the wrong tool?

A usability test tells you why people struggle; it does not tell you how much a change moves a business metric. If the question is "will this new checkout raise conversion?", you need live traffic and a control group, and usability testing vs A/B testing covers how to choose and sequence the two. If the question is whether a problem is worth solving at all, start with customer interviews instead.

The honest limitation of the 5-user rule is that it assumes one fairly uniform user group and a design with real problems to find. It is weak at spotting rare issues that affect a small share of users, and useless for measuring small differences between two good designs. For those, test more users or use a quantitative method.

How can you get better at running usability tests?

The Usability Testing for Product Managers bootcamp is a 1-week programme in the Discovery Expert Track, with 6 self-paced microlessons, directed by Andre Albuquerque. Its public curriculum covers test planning and goals, recruiting and screening, task design and scripts, moderation, synthesis and prioritisation, and turning findings into roadmap items. For further reading, the best usability testing resources for product managers collects articles and talks on the method.

Builders Camp runs live and self-paced bootcamps in product management and AI product building. See the Usability Testing for Product Managers bootcamp for the next cohort and the self-paced version.

Bootcamps referred in this Guide

Frequently asked questions

How do you run a usability test step by step?

Plan (name the decision the test feeds and the tasks that matter), recruit about 5 participants per distinct user group, write a session guide with realistic task scenarios, moderate with think-aloud and neutral prompts, rate each problem by frequency, impact and persistence, and report the top issues as decisions and roadmap items.

How many participants do you need for a usability test?

About 5 per round, per distinct user group. Jakob Nielsen of Nielsen Norman Group found that a first study with 5 users finds about 85% of the usability problems, and recommends several small rounds over one large study.

What should a PM own in a usability test if there is a researcher?

The decision the test feeds, the research questions, the choice of flows to test, observing at least some sessions live, and turning findings into roadmap items. The researcher usually owns the screener, the session guide wording, moderation and the analysis method.

How long does a usability test take to run?

A lightweight moderated test can run within a week: a day or two to plan and recruit, a day of sessions at around 30 to 60 minutes each, and a day to analyse and report. Recruiting niche participants is usually the slowest step.

What is the difference between a task and a question in a usability test?

A task asks the participant to do something with the product in a realistic scenario, such as booking a delivery for Saturday. A question asks for an opinion, such as whether they like the booking page. Tasks reveal behaviour; questions mostly reveal what people think you want to hear.

How do you rate the severity of usability problems?

Nielsen Norman Group combines three factors: how often the problem occurs, how hard it is to overcome, and whether it keeps bothering users after they learn about it. Its 0 to 4 scale runs from not a usability problem to a usability catastrophe that must be fixed before release.

Should a usability test be moderated or unmoderated?

Moderated when you need to understand why people struggle and can follow up in the moment, which suits early prototypes. Unmoderated when tasks are simple, you need more participants quickly, or you are checking that a fix worked.

Sources

Written by

Andre Albuquerque

Andre Albuquerque

CEO of Builders Camp, SuperOperator, and other companies. Building products.

CEO of Builders Camp, SuperOperator, and other companies. Building products.

LinkedInMore guides by Andre Albuquerque

Last updated 2026-09-27

Researched from Builders Camp's bootcamp, track and masterclass material and the sources listed on this page, drafted with AI, and fact-checked against every source cited.

See the Usability Testing bootcamp