Research MethodsTYPENORMLabs8 minJuly 29, 2026

Usability Testing — A Step-by-Step Guide

How to run a usability test end to end: framing the question, picking moderated or unmoderated, recruiting five people, writing tasks that don't lead, and turning what you watched into a ranked list of fixes.

The most expensive part of a usability test is the part nobody schedules: deciding what you're trying to find out. Teams book the participants, borrow a meeting room, and open the prototype — then spend an hour watching five people click around a product while the observers privately confirm whatever they already believed. The session happened. Nothing changed.

Usability testing answers exactly one kind of question: can a real person complete this task with this interface? Liking it, paying for it, wanting the next thing — different studies, different methods. Get the question right and the rest of the process is mechanical. Get it wrong and no amount of recruiting rigor saves the study. What follows is the sequence that produces a ranked list of fixes instead of a folder of recordings, plus the two habits that quietly invalidate most in-house sessions.

What usability testing measures (and what it doesn't)

A usability test puts a person in front of an interface, gives them a realistic goal, and watches what happens. You measure three things: whether they finished, how long it took and where they hesitated, and what they believed was happening at each step. Everything else in the room is noise.

The boundary matters because usability testing keeps getting asked to do jobs it can't. Whether anyone wants the feature is interviews and demand evidence. How common a problem is across your whole user base is analytics or a survey. What a test gives you is whether the path works, and it beats every other method at that because it's behavioral: you watch what someone does instead of collecting what they say they'd do. The gap between those two is where most product decisions go wrong. The wider map of which method answers which question lives on the UX research methods hub; this piece is the usability leg of it.

One consequence worth stating early: a test can only fail a design, never certify it. Five people completing checkout doesn't mean checkout works. It means you didn't find a blocker with those five people on that path. A clean run is an absence of evidence. Read it that way.

Step 1 — Write the question before the tasks

Start with a sentence that could turn out to be false. "Can a first-time user get from the pricing page to a completed signup without help?" is testable. "Is our onboarding good?" isn't: no observation would settle it, so the session drifts into opinion collection.

Write the question down and pin it where the observers can see it. It does two jobs. It decides which tasks belong in the script, and it gives you a stopping rule — when you can answer it, you're done, even if there's time left on the calendar invite.

If you have several worries, rank them and take the top two. A test that chases six questions answers none of them, because each additional task shortens the time on every other one and fatigues the participant right when you need their attention most.

Step 2 — Pick the method that fits the question

Two axes cover nearly every usability testing setup you'll consider.

Moderated or unmoderated. A moderator can ask "what were you expecting there?" the instant someone hesitates, and that follow-up is where the why comes from. Unmoderated sessions run through a tool that records screen and voice while the participant works alone: more volume, faster, far cheaper per person. Rule of thumb — moderated when you don't yet know what's broken, unmoderated when you have a specific hypothesis and want to watch it fail repeatedly.

Remote or in-person. Remote usability testing is the default now, and mostly that's fine. You get geographic range, cheaper recruiting, and people on their own hardware, which is more realistic than your loaner laptop anyway. In-person still wins when the physical context is part of the problem (kiosks, point-of-sale, anything held one-handed on a train) or when you need to see a face that a webcam flattens.

Fidelity is a third choice and a smaller one. A clickable prototype, a competitor's live product, a paper sketch, a staging build — all testable. Paper surfaces conceptual confusion; a live build surfaces flow and performance problems. Low fidelity is never a reason to postpone.

Step 3 — Recruit five of the right people

The famous number is five, and it holds up for one narrow claim: five participants from a single user group will surface roughly 85% of the usability problems on that path (NN/g's original analysis). Beyond five you start watching the same failures again, and the sixth session teaches you less than a second round on a fixed design would.

The load-bearing words are single user group. Five people means five people like each other in the way that matters for the task, because the maths behind the number assumes they're drawing from the same pool of possible confusions. Break that assumption and the 85% evaporates quietly, while the study still looks like it worked.

Most products break it. If yours serves both administrators and end users, that's two groups, two mental models, two failure shapes — and two sets of five. The admin knows what a "seat" is and fails on bulk actions; the end user has never seen the word and stalls on the invite screen. Run three of each and you've got a study that surfaces neither group's real problem while producing enough footage to feel thorough. That's the expensive kind of wrong: the deck gets built, the fixes get shipped, and the completion rate doesn't move.

Recruit on behavior. "Has switched analytics tools in the last year" predicts how someone will approach your import flow. "28–35, works in tech" predicts nothing at all — it's a description of your office.

Then screen out anyone who has seen the design before: a colleague, a friend of the PM, a beta user from three sprints ago. This is the first of the two habits that quietly invalidate most in-house sessions, and it's the one teams commit by accident, because the easiest person to book is the person who's already around.

Prior exposure is fatal and invisible. It doesn't look like contamination on the recording, it looks like a participant who breezed through checkout, which is the single most reassuring and least informative thing that can happen in a usability test.

Step 4 — Write tasks, not questions

A task is a goal with a reason attached and no instructions inside it. Compare:

  • Weak: "Click Settings, then Billing, and update the card." (You just did the test for them.)
  • Weak: "What do you think of the billing page?" (Collects an opinion. You came for a behavior.)
  • Strong: "Your company card was replaced last week. Make sure your next invoice charges the new one."

The strong version names a realistic motivation and leaves every navigation decision to the participant, which is the entire thing you're trying to observe. Scrub your own product's vocabulary out of the wording too. If the task says "open your workspace dashboard" and the nav item is labeled Workspace, you've handed over the answer and you'll never learn that nobody knows what that word means.

Order tasks the way a real session would unfold, front-load the one you care most about, and pilot the whole script on one person before the real sessions start. The pilot exists to catch tasks that are ambiguous, impossible, or already broken in the build — problems that would otherwise burn a real participant.

Step 5 — Run the session without leading

Make the participant comfortable, then get out of the way. Say plainly that you're testing the product and not them, and that confusion is the data you came for. People apologize for struggling, and every apology is a moment they've stopped narrating.

Then hold three lines. Ask for thinking aloud and prompt with a neutral "what are you thinking?" when they go quiet. Answer every question with a question ("what would you expect to happen?"). And let silence sit: the pause before someone clicks is the most informative second in the session, and moderators rush to fill it.

That middle one is the second habit, and unlike the first it's committed live, by the most senior person in the room, out of helpfulness. The moment you explain the interface you have deleted the finding. The participant now knows something no real user would know, and every observation after it belongs to your explanation rather than your design. Teams who fix their recruiting and not this still go home with a clean study that means nothing.

Observers take timestamped notes in a shared doc, keeping what happened separate from what it might mean. That split makes the next step fast instead of argumentative.

Step 6 — Turn observations into a ranked list

After the last session, pull every observation into one place and cluster the ones describing the same underlying breakage. Three people hesitating at the same unlabeled icon is one issue with three instances.

Rank each cluster by severity on two inputs: how badly it blocked the task (annoyance → workaround → hard stop) and how many participants hit it. A hard stop one person hit outranks a mild irritation all five mentioned, because the stop costs you the conversion and the irritation costs you a moment. Write each issue as evidence plus a proposed change — "4/5 missed the plan selector because it reads as a static heading; make it a labeled control" — so the fix is arguable on its merits rather than on whose intuition is louder.

Ship the top few, then test the same path again. Usability testing pays off across rounds: a small study on a design you changed last week beats a large study on a design nobody has touched.

Usability testing vs heuristic evaluation

The two get confused because both produce a list of problems. They're different instruments.

Usability testing observes real people and finds problems you couldn't have predicted: the ones that come from mental models you don't share. Slower, more expensive per finding, and the only method that tells you what actually happens.

Heuristic evaluation has two or three experts review the interface against established principles (Nielsen's ten, for instance) and flag violations. Fast, cheap, and it catches a broad sweep of known issues in an afternoon. What it produces are predictions, and experts reliably miss the confusions that come from not being an expert.

The efficient order is heuristic evaluation first, to clear the obvious breakage, then usability testing on the paths that survive — so you don't spend a participant's hour discovering a missing error message a reviewer would have caught for free. That's the same sequencing behind our Full UX Audit, which starts with expert review precisely so testing time gets spent on the questions review can't answer.

Frequently asked questions

How many users do you need for a usability test? Five per user group, per round. Quantitative usability testing is a different study: measuring completion rates or task times at statistical confidence needs 20+ per group.

How long should a session be? Under an hour. Past that, fatigue starts producing failures that belong to the participant's attention span and not to your interface.

Can you run usability testing without a budget? Yes. Five people, a video call with screen sharing, and a clickable prototype is a legitimate study. Everything that actually ruins a test is free to fix: the vague question, the leading task, the participant who already knows the product.

When in the process should you test? Whenever a decision is expensive to reverse. Early on that's a paper concept; mid-build it's a prototype; post-launch it's the live flow. The one guaranteed-wrong time is right before release, when nothing surfaced can actually be changed.

Is usability testing the same as A/B testing? No, and the two are complements rather than rivals. A/B testing takes a variant you've already built, shows it to a lot of traffic, and reports which one performed better. It's the only honest answer to "how much is this worth," and it is silent on why. Usability testing takes five people and tells you why, in detail, with no ability to say how widely the problem holds.

So the sequence writes itself. Testing generates the hypothesis: five people all miss the plan selector, and here's what they thought it was instead. A/B then prices it, by shipping the labeled control to half your traffic and reading the difference. Teams that run only A/B end up optimizing variants nobody understands; teams that run only usability testing end up confidently fixing things that were never costing them anything.

The short version

Usability testing is cheap and forgiving in every way except one: it can't recover from a badly framed question.

Everything else on this page is recoverable. A weak task gets rewritten for the next participant, a bad recruit gets replaced, a moderator who talks too much learns by the third session. But the question you wrote at the top decides what the whole study is capable of finding, and you only get to choose it once.

Then do it again in two weeks on the thing you fixed. The compounding is the point.

Free UX Snapshot for 50 Product Teams

Apply now and get a complimentary UX Snapshot — our rapid clarity audit delivered in 48 hours. Limited to the first 50 products.

Apply for Free UX Snapshot

Related

Information Architecture

UX Collective UX Teardown: Building the Biggest Design Publication on Rented Land

A UX teardown of UX Collective: the largest independent design publication runs entirely on Medium, and every page is shaped by owning the taste but not the platform. How the read flow races Medium for the reader — and the one tab that wins.

TYPENORMLabs · 5 min · July 24, 2026

Web
Information Architecture
Interaction Design

Accessibility

Skip Navigation Links for Accessibility

A skip navigation link lets a keyboard user jump past the nav straight to the content. Four lines of markup, and almost every implementation gets the CSS or the focus wrong.

TYPENORMLabs · 7 min · July 25, 2026

Accessibility
Web

Information Architecture

Nielsen Norman Group UX Teardown: The Oldest Funnel in UX, Run in Plain Sight

A UX teardown of Nielsen Norman Group's site: how the firm that made usability a profession turns a free article library into five- and six-figure consulting — three tiers of the same authority, priced for three different buyers.

TYPENORMLabs · 5 min · July 24, 2026

Web
Information Architecture