Skip to content
Back to the journal

Essay

047

Discovery & Validation

12 min read

047 / 136

User Interviews: Design Evidence You Can Challenge

Plan, conduct, and analyse user interviews as bounded evidence: connect the decision, sample, questions, observations, interpretation, and limitations.

Updated July 13, 2026

Topics User research Validation Discovery

Share this essay

A convincing interview can be worse than an inconclusive one.

The participant speaks fluently about the problem. Their language matches the team’s hypothesis. A memorable quote reaches the roadmap deck by the end of the day.

Nobody asks whether the participant recently faced the situation, whether the interviewer supplied the framing, or whether a contradictory account was discarded as an outlier.

The conversation was rich. The evidence chain was weak.

A user interview is not direct access to truth. It is an account produced by a particular participant, with a particular interviewer, under particular conditions, for a particular research purpose.

Designing a good interview means preserving that chain from decision to claim. The goal is not to eliminate interpretation. It is to make the interpretation visible enough to challenge.

Start with the decision that could change

“Understand users better” cannot determine whom to recruit, what to ask, or when the work is sufficient.

Name the decision first:

  • Which product choice is open?
  • What does the team currently believe?
  • Which alternative explanations remain credible?
  • What evidence would change the choice, narrow it, or stop it?
  • What is outside the interview’s authority?

Suppose a team sees fewer small-business administrators complete account setup. One explanation is that the workflow is confusing.

Others remain possible: administrators may lack required information, need a colleague’s authority, distrust the request, or decide the product is not worth the effort.

“Why do users abandon onboarding?” is too broad. A more useful research question might be:

How do first-time administrators decide whether, when, and with whose help to complete the required setup work?

That question invites accounts of a decision and its context. It does not ask participants to endorse the team’s preferred cause.

The GOV.UK Service Manual says research questions can help plan and prioritise research, and distinguishes those questions from the words asked in a session.

Its research-question guidance is written for government services.

The transferable point is that a learning objective and an interview prompt do different jobs.

Check whether an interview can answer it

Interviews are useful for learning how people describe situations, meanings, decisions, constraints, relationships, and remembered experience.

They are weaker tools for several other claims.

  • To see whether somebody can complete a task with an interface, observe the attempt in a usability study.
  • To estimate prevalence in a population, use an appropriate quantitative design and sampling method.
  • To know what the product recorded, inspect telemetry and its measurement contract.
  • To understand work that is difficult to articulate, observe the work or examine its artefacts.
  • To estimate causal effect, use a design capable of supporting that inference.

An interview can help explain an observed pattern or generate a hypothesis. It cannot turn a small, purposive sample into a population percentage.

It also cannot make a future promise behave like past action. “Would you use this?” may reveal a reaction to the concept and the language used to present it.

It does not establish adoption under real cost, competing priorities, organisational permission, or repeated use.

Choose the method from the claim. If a mixed design is needed, decide how its parts will disagree productively rather than using one method as decoration for another.

Recruit for the variation the decision needs

The right participant is not simply “a user.” It is someone whose relationship to the decision makes their account relevant.

Build a sampling frame from the important variation:

  • role, responsibility, and decision authority;
  • recent experience of the situation;
  • product or workaround currently used;
  • success, failure, abandonment, and non-use;
  • frequency or maturity of the work;
  • relevant access needs and support relationships;
  • market, language, policy, or operating context.

Recruiting only engaged customers makes access easy and may remove the people whose barriers matter most. Recruiting only buyers can hide the work of operators.

Recruiting whoever answers fastest can turn availability into an accidental product segment.

Do not pretend a purposive sample is statistically representative. Instead, state which perspectives it was designed to include, which it missed, and why those omissions matter.

GOV.UK’s participant guidance explicitly notes that unbiased recruitment is difficult.

It recommends varying recruitment approaches and including disabled participants and people with access needs. Those requirements apply directly to UK government service teams.

For other product organisations, the broader discipline is still useful: examine how the recruitment channel, schedule, location, incentive, eligibility screen, and research format include some experiences and exclude others.

Build a route through evidence, not a question script

A discussion guide should create consistency without closing the conversation.

Organise it around evidence you need, not every question somebody on the team wants answered.

Establish the situation

Ask the participant to locate a concrete episode: when it happened, what triggered it, what they were responsible for, who else was involved, and what counted as a satisfactory outcome.

“Tell me about the last time” is useful when the latest episode is relevant and recall is credible. It is not a law.

Atypical recent events, rare high-consequence decisions, or sensitive topics may require a different anchor.

Reconstruct the sequence

Ask what happened next. Explore tools, hand-offs, waits, workarounds, interruptions, and decisions. Invite the participant to show a redacted artefact or workflow where consent and confidentiality allow.

Separate what the participant did from what they usually do, what policy says should happen, and what they wish would happen.

All four can matter. They are not the same evidence.

Probe the consequence

What changed because of the action? Who absorbed the cost? How did the participant know the work was complete or acceptable?

Words such as “difficult,” “manual,” and “important” are openings, not findings. Ask what made the situation difficult, what the manual work consisted of, and what happened when it was not done.

Search for the counterexample

Ask when the problem does not occur, who handles it differently, or what made a similar situation easier.

A counterexample can reveal a missing condition more effectively than another confirming account.

Separate concept reaction

If the session includes a concept, mark the transition. First preserve evidence about the participant’s world; then expose them to the team’s artefact.

Record exactly what they saw and in which state. A reaction to an explanation, static screen, and functioning workflow are not interchangeable.

The GOV.UK guide to in-depth interviews recommends open, neutral prompts, real examples, a tested discussion guide, and flexible follow-up.

It is public-service practice guidance rather than comparative evidence that one guide format is superior. Use it as a sound operating reference, not a universal protocol.

Do not manufacture the account

Interview technique affects the evidence. The interviewer chooses the topic, asks the next question, signals interest, and decides when an answer is complete.

That influence cannot be removed. It can be governed.

Use neutral invitations before interpretations:

  • “What happened then?”
  • “How did you decide?”
  • “You used the word ‘blocked’. What did that mean in this situation?”
  • “What else was different that time?”
  • “Can you show me, if it is safe to do so?”

Avoid offering the desired answer inside the question. “Was that frustrating because approvals are slow?” supplies an emotion and a cause.

“How did the approval affect the work?” leaves both open.

Do not repeat “why?” as an interrogation ritual. It can invite post-hoc explanations and imply that the previous answer was inadequate.

Move between sequence, evidence, consequence, and reflection. Ask for clarification without demanding a neat root cause.

Silence is not automatically depth. Body language is not a hidden truth detector. If hesitation seems consequential, ask rather than diagnose it.

The participant controls what they disclose. Treat consent as an ongoing boundary, not a form collected before the interesting part begins.

State who is observing, what is recorded, how material will be used, and what the participant may decline. Stop or redirect when the conversation crosses the agreed purpose.

Keep observation and interpretation separable

During and immediately after the session, preserve several layers:

  1. Context: participant characteristics relevant to the research question and the conditions of the session.
  2. Account: what the participant said, with quotations only where the wording matters.
  3. Observed action: what they did or showed, distinct from their explanation of it.
  4. Researcher interpretation: what the team thinks the evidence may mean.
  5. Alternative interpretation: another credible account of the same material.
  6. Open question: what remains unresolved or requires another method.

Do not make the quotation carry a conclusion it cannot support. A quote can illustrate a theme and preserve language.

It cannot prove that the theme is common, important to every user, or causally responsible for a metric movement.

Research notes and recordings can contain personal and commercially sensitive data.

GOV.UK’s data and privacy guidance requires informed consent and purpose-bounded use for government research.

It also requires limited collection, controlled access, secure storage, and timely deletion.

Other organisations must apply the law and policy relevant to them. At minimum, a research repository should not become a permanent archive merely because storage is cheap.

Turn interviews into claims with boundaries

Analysis is not the moment when “themes emerge” by themselves. Researchers select, compare, interpret, and name patterns.

For product decisions, a claim-evidence ledger keeps that work inspectable:

Claim
Relevant participants and situations
Supporting accounts or observations
Contradicting or absent evidence
Researcher interpretation
Alternative explanations
Sample and method limitations
Decision implication
Confidence and next evidence

Start within cases before collapsing across them. Reconstruct each participant’s situation, sequence, and consequence.

Then compare cases on the dimensions that matter to the decision. A buyer, daily operator, failed adopter, and administrator may describe different systems rather than disagree about one system.

Count participants only when the count has a legitimate role in the analysis. “Six of eight mentioned permissions” describes this sample.

It does not estimate population prevalence. Frequency can also hide consequence: a rare account may expose a serious safety, exclusion, or operating risk.

Look for negative cases. If the claim is “approval complexity causes abandonment,” examine participants who faced complexity and still completed, as well as those who abandoned without it.

The purpose is not to make a qualitative study mimic an experiment. It is to prevent the first coherent story from becoming the only story.

A fictional decision trail

Consider an explicitly fictional B2B platform whose activation data show that some new workspace administrators do not invite a second member.

The team does not begin by asking, “Would collaboration features make onboarding better?” It opens a decision: whether to invest in an invitation redesign, change the setup sequence, or investigate a different constraint.

The sample includes recent completers, recent non-completers, administrators who needed procurement approval, and administrators who work alone. The team records that it did not reach organisations with strict identity policies.

Interviews reconstruct the most recent setup attempt. Some participants describe invitation copy; others reveal that they were evaluating the product privately and had no authority to involve colleagues.

One participant completed the invitation only after a security review. Another did not invite anyone because the workspace represented a single-person client account.

The team writes a bounded claim: the invitation step combines several materially different situations that the current event cannot distinguish.

It does not claim that invitations cause abandonment or that a redesigned screen will improve activation.

The next decision is to separate eligibility and intent in measurement, then test the relevant interaction for administrators who are ready and authorised to invite.

No result is invented. The value lies in the evidence chain: metric pattern, open decision, sampled variation, concrete accounts, contradictory cases, limitation, and next method.

Publish the research disposition

A research readout should end with more than findings.

For each material claim, record whether the team will:

  • act within the supported boundary;
  • seek discriminating evidence;
  • change the product question;
  • preserve a known disagreement;
  • decline to act because the evidence is too weak;
  • revisit the claim when a named condition changes.

Link the decision back to the evidence and the evidence to its collection conditions. Later teams should be able to see what was known, not just inherit a quote stripped of context.

Continuous Research explains how to govern the cadence and decision queue around repeated studies.

Hypothesis Testing is the adjacent practice when the team must make competing explanations answerable to discriminating evidence.

An interview has done useful product work when it changes the quality of a decision and leaves the limits of that change visible.

Sources

Related books

If you want to go further on this topic, these are two good places to start.

02

product

The Lean Startup

by Eric Ries

How today's entrepreneurs use continuous innovation to create radically successful businesses, introducing Build-Measure-Learn and validated learning.

Some outbound links are affiliate links and support independent bookstores.