Usability Studies: Test a Use Risk, Not the Interface
Design usability studies around specified users, goals, contexts, and release decisions—then turn observed breakdowns into bounded findings and action.
Piotr Ciechowicz
Product manager · developer
Updated July 13, 2026
On this page15 sections
- 01Define the usability claim
- 02Choose discovery or verification deliberately
- 03Formative study: find and explain breakdowns
- 04Comparative study: choose between credible alternatives
- 05Summative study: assess a requirement
- 06Recruit for variation that can change the result
- 07Make tasks credible without giving away the route
- 08Fix the protocol before watching the first person
- 09Observe events before explaining them
- 10Treat evaluator variation as evidence risk
- 11Turn an issue into a decision
- 12A fictional study that changes the release question
- 13Close with a bounded claim
- 14Sources
- 15Read next
“Test the new flow” sounds specific until the results arrive. One participant misses a control, another dislikes the wording, and a third completes the task after the moderator helps.
The team has observations but no rule for what they mean.
A usability study should examine a bounded risk in use. It needs specified users, a goal, a context, observable conditions, and a decision that the evidence can change.
Without those boundaries, the session becomes a tour of the interface. A vivid struggle can dominate the room, an easy completion can be treated as proof, and neither supports a responsible release decision.
Define the usability claim
ISO 9241-11:2018 treats usability as an outcome of use in relation to specified users, goals, and context, considered through effectiveness, efficiency, and satisfaction.
The standard defines concepts; it does not prescribe a testing method.
That definition prevents “the product is usable” from functioning as a universal claim. A dispatch tool may support a trained operator on a desktop while failing a substitute operator on a tablet during an outage.
Start the study brief with six fields:
| Field | What to specify |
|---|---|
| Decision | The release, redesign, comparison, or investigation this evidence will inform |
| Population | Relevant capability, experience, access need, role, and exclusion—not just a demographic label |
| Goal and tasks | What participants must achieve, including recovery or exceptional work |
| Context | Device, environment, information, pressure, assistance, and connected systems |
| Use risk | The breakdown and consequence the team needs to understand |
| Claim boundary | What this study can and cannot support |
NISTIR 7432 uses a similar structure for usability requirements: context of use, performance and satisfaction criteria, and the test method and context used to assess them.
The NIST document is a requirements specification, not evidence that one template improves products.
Its value here is traceability: a result is interpretable only against the use condition and criterion it was designed to examine.
Choose discovery or verification deliberately
Formative and summative work can both involve people completing tasks, but they carry different claims.
Formative study: find and explain breakdowns
Use formative work while a design can still change. The study may ask where people form the wrong expectation, which state is invisible, why recovery fails, or which wording conflicts with domain practice.
The output is a set of bounded findings and design decisions. Task completion and time can help locate an event, but a small purposive sample should not be turned into a population success rate.
Comparative study: choose between credible alternatives
Use a comparison when the team has genuine alternatives and a decision rule. Control task wording, order, data, and moderator behaviour enough to make the difference interpretable.
If one prototype has complete data and the other has placeholder content, the study compares more than interaction design. Record those differences instead of attributing the result to whichever element the team hoped to test.
Summative study: assess a requirement
Use summative testing when the decision depends on whether defined criteria are met under a specified protocol.
This requires an appropriate sampling plan, operational measures, consistent administration, and analysis agreed before results are seen.
A few exploratory sessions can discover consequential problems. They cannot certify compliance, establish a stable benchmark, or estimate population performance merely because the team calculated percentages.
Recruit for variation that can change the result
“Five users” is not a recruitment strategy.
Faulkner tested 60 people on one web-based employee timesheet application, then repeatedly sampled subsets. Five-person groups found between 55% and 99% of the problems known in the full dataset.
The minimum rose to 82% for ten-person groups and 95% for twenty.
That 2003 study demonstrates variability in one product and participant pool. It does not provide a universal minimum or guarantee that 20 participants reveal 95% of problems in another study.
Set the sample from the decision and the variation that matters:
- distinct roles or permission levels;
- novice, intermittent, and experienced use;
- assistive technology and access needs;
- device, channel, language, or environmental conditions;
- common paths and low-frequency, high-consequence tasks;
- prior exposure that changes terminology or mental models.
Use a coverage matrix. Rows represent participant characteristics that can plausibly change use; columns represent tasks and contexts. Empty cells show what the study does not cover.
Do not recruit “wrong users for a fresh perspective.” A person outside the intended or affected population can evaluate a general interaction convention.
Their behaviour cannot stand in for expertise, obligations, incentives, or access needs they do not have.
If disabled people are in scope, plan the environment, consent, assistive technology, communication, and support with them. Usability observation complements but does not replace standards-based accessibility evaluation.
The inclusive-design guide explains how to make exclusion and exceptions explicit product decisions.
Make tasks credible without giving away the route
A good task states a goal in language the participant could plausibly encounter. It supplies the information they would have and withholds the interface label the study is trying to test.
“Use Advanced Filters to find unpaid invoices” tests whether a person can follow an instruction. “Your finance lead needs the invoices still unpaid after 30 days” leaves the route open.
For each task, define:
- starting state, data, permissions, and prior knowledge;
- observable success and consequential partial success;
- critical error or unsafe state;
- time boundary only when time matters to the real goal;
- assistance the moderator may give and how it will be recorded;
- the state to restore before the next task.
GOV.UK guidance similarly recommends clear, believable tasks that answer the research questions without hinting at the solution. It is practice guidance for public services, not a controlled comparison of task-writing techniques.
Prototype fidelity should follow the risk. A clickable model may be sufficient to test navigation or wording.
It is weak evidence for latency, permission behaviour, generated content quality, cross-device recovery, or a workflow that depends on real records.
Use the rapid-prototyping guide to state what the artefact simulates, implements, and excludes.
Fix the protocol before watching the first person
A protocol protects the claim from improvisation. It does not require the moderator to behave like a machine.
Record:
- the introduction, consent, and recording agreement;
- the neutral wording and order of tasks;
- what the moderator says when a participant pauses, asks for help, or goes off path;
- when a task stops and how assistance is coded;
- which follow-up questions wait until after the observed attempt;
- what observers capture and how identifying data is protected;
- changes made to the prototype or protocol between sessions.
Run a pilot. A pilot can reveal missing data, impossible starting states, leading prompts, broken recording, or a task whose wording is harder than the product.
Do not quietly repair the protocol after one participant and combine all sessions as though conditions were unchanged. Improve it when needed, version the study, and preserve which evidence came from which condition.
GOV.UK advises telling participants that the service, not the person, is being tested; using neutral instructions; mostly observing; and recording consent and personal data securely.
Those are ethical and practical safeguards, not proof of a particular study’s validity.
Observe events before explaining them
Separate the session record into three layers:
- Event: what was shown, said, selected, omitted, or requested.
- Interpretation: the plausible expectation, constraint, or mechanism behind it.
- Finding: a bounded claim supported across the relevant evidence, with counterevidence and limits.
“Participant 4 selected Download before choosing a reporting period” is an event. “Users expect the period after export” is an interpretation until evidence supports a broader statement.
Record assistance as an event, not an invisible moderation success. A participant who completes after a hint has shown both a recovery path and a breakdown under unaided use.
Use recordings only within the consent given. A highlight reel can help a team see an event, but editing removes context and amplifies memorable moments. It should link back to the session record, not replace analysis.
For broader standards on claim construction, counterevidence, and limits, use Qualitative Product Research.
Treat evaluator variation as evidence risk
Analysis is not a neutral extraction process.
In CUE-4, 17 experienced professional teams independently evaluated the same hotel website. Nine used usability tests and eight used expert reviews. Together they reported 340 different issues.
Only nine issues appeared in reports from more than half the teams, while 205 issues were unique to one team. Sixty-one of those unique issues were classified serious or critical.
The study does not reveal which team found “the truth,” and it did not establish missed problems or false alarms for expert reviews. It shows that professional teams can produce sharply different issue sets from the same site.
Reduce that risk by:
- having at least two people challenge material findings;
- tracing every finding to session events;
- actively looking for successful and contradictory cases;
- keeping task, participant, context, and prototype differences visible;
- separating observed consequence from anticipated production consequence;
- recording unresolved interpretations rather than forcing consensus.
More observers do not automatically create reliability. Shared assumptions can produce shared blind spots.
Turn an issue into a decision
An issue count is not a severity model. Frequency in a small study is not population prevalence, and one occurrence can still expose a severe failure mode.
Create one record for each material finding:
Affected goal and population
Observed events and counterevidence
Plausible mechanism
Consequence in this study
Potential production consequence and uncertainty
Current evidence boundary
Proposed disposition and owner
Retest condition
Use one of five dispositions:
- change now because the evidence exposes a violation or unacceptable use risk;
- investigate because the mechanism or reach is unclear;
- compare because credible design alternatives remain;
- accept temporarily with an owner and reopen trigger;
- decline because the issue is unsupported, outside the product boundary, or outweighed by stronger evidence and constraints.
A usability study does not decide product value or demand. A design can be easy to use and solve no important problem. That adjacent decision belongs to concept validation.
A fictional study that changes the release question
Consider a fictional logistics product replacing email approval for delivery exceptions. The team wants to know whether dispatchers can resolve an exception before the vehicle leaves the depot.
The original plan tests a polished approval screen with experienced dispatchers. The study brief exposes a missing use condition: a substitute dispatcher may inherit the queue after the evidence was collected by somebody else.
The team recruits for both roles and adds a recovery task. The prototype contains plausible timestamps, evidence states, and permissions.
The protocol defines completion without help, completion after a neutral prompt, and an unsafe approval without required evidence.
In this fictional example, an observed breakdown would not prove its prevalence. It could still block release if the team had already defined evidence-free approval as an unacceptable consequence.
The decision is no longer “Do people like the new screen?” It is whether the workflow preserves enough state and authority for the intended population to resolve the exception safely.
Close with a bounded claim
A useful readout states:
- what was studied and why;
- who and what context the evidence represents;
- which tasks and conditions were covered;
- what broke, what succeeded, and what remains disputed;
- which inferences go beyond direct observation;
- what decision changed and who owns it;
- what the study cannot support.
The strongest usability report is not the one with the most findings. It is the one that lets another person inspect how an observed event became a product decision—and where that reasoning should stop.
Sources
- International Organization for Standardization. ISO 9241-11:2018 — Usability: Definitions and concepts. The current standard defines usability and its application; it does not prescribe evaluation methods.
- Theofanos, M. F. (2007). Common Industry Specification for Usability — Requirements, NISTIR 7432. A specification for context, criteria, test method, and traceability, not an outcome study.
- Faulkner, L. (2003). Beyond the five-user assumption: Benefits of increased sample sizes in usability testing. Behavior Research Methods, Instruments, & Computers, 35(3), 379–383.
- Molich, R., & Dumas, J. S. (2008). Comparative usability evaluation (CUE-4). Behaviour & Information Technology, 27(3), 263–281.
- GOV.UK Service Manual. Using moderated usability testing. Public-service practice guidance on tasks and moderation.
- GOV.UK Service Manual. Taking notes and recording user research sessions. Public-service practice guidance on consent, privacy, and records.
Read next
Related books
Two books to
read next.
If you want to go further on this topic, these are two good places to start.
01
product
Continuous Discovery Habits
by Teresa Torres
A practical guide to discovering products that create customer value and business value, with frameworks for integrating customer research into weekly rhythms.
02
product
The Lean Startup
by Eric Ries
How today's entrepreneurs use continuous innovation to create radically successful businesses, introducing Build-Measure-Learn and validated learning.
Some outbound links are affiliate links and support independent bookstores.