Product Discovery Techniques: Build an Evidence Portfolio
Choose product discovery methods by the uncertainty they can reduce, the decision they support, and the claims their evidence cannot justify.
Piotr Ciechowicz
Product manager · developer
Updated July 13, 2026
On this page15 sections
- 01Make uncertainty the unit of planning
- 02Give each kind of uncertainty a different evidence job
- 03Context: what is actually happening?
- 04Distribution: where and how often does the pattern appear?
- 05Interaction: can people understand and use the response?
- 06Feasibility: can the system deliver the promise?
- 07Commercial and organisational fit: can the value be adopted and sustained?
- 08Impact: did the product change the outcome?
- 09Build combinations, not a procession of methods
- 10Match the evidence burden to the commitment
- 11A hypothetical portfolio for an offline workflow
- 12Watch for method debt
- 13Keep a one-page evidence portfolio
- 14Sources
- 15Read next
Many discovery plans are calendars in disguise.
Interview customers on Monday. Map the journey on Wednesday. Test a prototype next week. The activities sound responsible, but nobody has explained why those methods belong together.
This is how a team collects plenty of evidence and still carries its most dangerous assumption into delivery.
A discovery method is useful only in relation to an uncertainty. Interviews can explain a workflow but cannot estimate its prevalence. A prototype can expose confusion but cannot prove that the underlying problem matters.
Treat discovery as an evidence portfolio. Combine methods with different strengths, reject evidence that cannot support the claim, and spend more effort where being wrong is expensive.
Make uncertainty the unit of planning
Start with one decision and the beliefs it depends on. Then inspect the evidence already available.
A compact opening statement is enough:
We are deciding whether to [commitment]. The decision depends on [belief]. We currently know [evidence], but remain uncertain about [gap].
The nine product discovery questions provide a fuller frame when the decision itself is still vague.
Do not name a method in the opening statement. “We need interviews” closes the choice too early. The real gap may concern distribution, technical behaviour, procurement, or causal impact.
GOV.UK’s research-planning guidance follows the same order: agree the research questions, identify the relevant users, and then choose activities that can answer those questions reliably.
The portfolio begins when a team can say what each method must learn and what it will leave unresolved.
Give each kind of uncertainty a different evidence job
The following portfolio is a practical synthesis, not a universal discovery taxonomy. Its purpose is to stop one familiar method from becoming the answer to every product question.
Context: what is actually happening?
Use contextual observation, interviews about recent events, support material, service records, or existing research to reconstruct the situation.
Look for triggers, sequence, participants, workarounds, constraints, and the outcome people protect. The result should describe behaviour in context rather than a list of requested features.
GOV.UK’s discovery guidance recommends examining the current experience end to end, including tools, support, transactions, and offline steps.
This evidence can explain how a problem appears. It cannot establish how common the pattern is across a population.
Distribution: where and how often does the pattern appear?
Use product telemetry, support classification, operational data, market records, or a carefully designed survey to examine the reach and variation of a pattern.
Define the denominator before interpreting the result. A large count may represent a common event in a large population or a severe concentration in one small segment.
Check whether the data could observe the behaviour at all. A failed offline action may never reach the event pipeline. A support taxonomy may hide one failure under several labels.
Distribution evidence can show concentration, trend, and variation. It rarely explains the mechanism by itself.
When the decision concerns scale, Research-Driven Opportunity Sizing helps turn these inputs into ranges and traceable assumptions.
Interaction: can people understand and use the response?
Use a prototype or working service with realistic tasks when uncertainty concerns comprehension, navigation, control, or recovery.
Moderated usability testing observes people attempting specific tasks. GOV.UK positions it as a way to identify task and interface problems in prototypes or existing services.
Test the risky state alongside the polished path. Include missing information, errors, permission limits, and the moment someone changes their mind.
A successful task supports a claim about usability under the tested conditions. It does not demonstrate demand, commercial viability, or durable behaviour after launch.
When a concrete response is already selected and the question is what the evidence permits the team to commit next, move into concept validation rather than extending the method catalogue.
Feasibility: can the system deliver the promise?
Use engineering spikes, data profiling, benchmarks, replay against representative cases, operational walkthroughs, or observation mode.
Match the evidence to the promise. “Instant” raises latency and load questions. “Always current” raises source-of-truth and delay questions. “Automatic” raises exception and recovery questions.
A technical demonstration is weak when it uses clean inputs that production cannot provide. Record data gaps, failure modes, operating cost, and the conditions under which the system must abstain.
Feasibility evidence can show that a response can be built and operated within stated constraints. It cannot show that the response deserves to exist.
Commercial and organisational fit: can the value be adopted and sustained?
Investigate who benefits, who approves, who pays, who carries risk, and which workflow must change. The user, buyer, administrator, and accountable owner may be different people.
Useful methods include buyer-process interviews, procurement and policy review, pricing research, a manual service, or a bounded commercial pilot.
Verbal enthusiasm is weak evidence when the decision requires budget, data access, migration, or a policy exception. Ask what the organisation has done in comparable purchases and which obstacle could stop action.
This work can expose adoption and operating constraints. It should not be turned into a universal willingness-to-pay score.
Impact: did the product change the outcome?
Use an experiment or a well-designed quasi-experimental evaluation when the decision depends on whether the intervention caused a change.
The 2026 Magenta Book explains that experimental and quasi-experimental approaches estimate impact by comparing outcomes with a credible counterfactual.
That requires comparable groups or periods, suitable data, and an effect distinguishable from expected noise. The evaluation design often needs to be planned before release.
Impact evidence can estimate whether a change caused an outcome under defined conditions. It may still leave the mechanism, affected subgroups, and wider context unclear.
Qualitative follow-up and segmented analysis can investigate those remaining questions without pretending that one aggregate result explains every experience.
Build combinations, not a procession of methods
More methods do not automatically create more confidence. Two methods may repeat the same blind spot.
An interview followed by a survey of people recruited from the same customer panel may broaden the format without broadening the evidence. A prototype test and a preference poll may both measure reaction rather than behaviour.
Combine methods when each closes a different gap:
| Starting uncertainty | First evidence move | Complementary move | Decision gained |
|---|---|---|---|
| The problem is poorly understood | Observe recent work and interview around specific events | Examine telemetry or operational records | Decide which pattern deserves sizing |
| The interaction may fail | Test the risky task in a prototype | Run a bounded release with recovery signals | Decide whether the design is ready to expand |
| The system promise may be infeasible | Profile data and run a technical spike | Observe outputs on representative live cases without acting on them | Decide whether to narrow, invest, or stop |
| Adoption depends on several roles | Map the user, buyer, administrator, and owner | Test a manual or limited commercial path | Decide whether the route to adoption is credible |
| A released change appears successful | Analyse the observed outcome and competing explanations | Design a suitable impact evaluation | Decide whether the change caused enough value to scale |
The sequence also matters. Testing a solution before understanding the situation can produce precise evidence about the wrong response.
Some uncertainties can run in parallel. Technical data access and buyer approval may be independent enough to investigate at the same time.
Others have dependencies. A causal experiment needs a stable intervention and reliable outcome measurement. A pricing test needs a defined buyer and value mechanism.
Assumption mapping helps expose those dependencies before the team fills a sprint with disconnected research.
Match the evidence burden to the commitment
A sketch used to choose between two conversation directions does not need launch-grade evidence. A migration affecting regulated records does.
Set an evidence burden by asking:
- How costly is reversal?
- Who could be harmed by a wrong decision?
- Which groups may experience the product differently?
- How much operational or commercial exposure will the commitment create?
- Which uncertainty could remain hidden until scale?
This does not produce a confidence score. It determines where the portfolio needs stronger coverage, independent evidence, or a narrower release.
Evidence quality has several dimensions:
- relevance: does it concern the actual people, task, and decision?
- coverage: which contexts or groups are missing?
- recency: does it describe the current product and environment?
- traceability: can someone inspect how the finding was produced?
- independence: does another method or source challenge the same claim from a different angle?
A large dataset can be irrelevant. A small observation can be decisive when it reveals that a required workflow is impossible.
A hypothetical portfolio for an offline workflow
Consider a fictional field-service product exploring offline approval of maintenance work.
The initial request is to “add offline mode”. That phrase hides several uncertainties.
The context stream examines recent jobs, device conditions, handovers, and the records technicians use when connectivity disappears.
The distribution stream checks which workflows lose connectivity, whether missing events distort the telemetry, and which roles encounter the problem.
The interaction stream tests conflict, stale information, queued actions, and reconnection. A happy-path prototype would miss the states that make offline work difficult.
The feasibility stream profiles local storage, synchronisation, permissions, and recovery. The commercial stream examines which customers can permit local data and who accepts responsibility for a delayed approval.
An impact test would come later. The team first needs a bounded, operable intervention and a measurable outcome such as completed approvals with correction and rework visible.
The portfolio does not return one “validated” score. It shows which claims became stronger, which remain exposed, and whether the next commitment should be a technical investment, a narrower pilot, or no offline product at all.
Watch for method debt
A team can become dependent on methods that fit its access rather than its decisions.
Interview monopoly: every question becomes a conversation, including questions about prevalence or system performance.
Analytics without meaning: the team sees a behavioural pattern but invents the explanation from a dashboard.
Prototype inflation: task completion is presented as evidence of demand or retention.
Experiment theatre: a test launches without a stable hypothesis, trustworthy instrumentation, or a decision attached to the result.
Evidence recycling: one old study appears in several presentations until repetition makes it feel current and universal.
Discovery without expiry: the team keeps learning because nobody defined what evidence would be sufficient for the next commitment.
Review the portfolio, not the volume of research. Repeated blind spots are a capability problem. The team may need different access, instrumentation, expertise, or a partner from another discipline.
Keep a one-page evidence portfolio
For each active product decision, record:
- the commitment and decision owner;
- the claims that must hold;
- current evidence and its limits;
- the most consequential uncertainty;
- the method selected and why it fits;
- what the method cannot establish;
- the result that would narrow, stop, or advance the work;
- the next review date.
Update the record when evidence arrives. Remove methods that no longer serve a live uncertainty.
Discovery becomes strategic when method choice reveals the shape of the bet. The goal is not a balanced research calendar. It is enough relevant evidence to make the next commitment with the risks visible.
Sources
- Plan user research for your service — GOV.UK Service Manual
- User research in discovery — GOV.UK Service Manual
- Using moderated usability testing — GOV.UK Service Manual
- Magenta Book: Central Government guidance on evaluation — HM Treasury
- Magenta Book Annex A: analytical methods for evaluation — HM Treasury
Read next
9 Product Discovery Questions That Force a Decision establishes the decision frame before the team selects evidence.
How to Approach Assumption Mapping Systematically helps identify which uncertainty should control the next research investment.
Discovery Habits: Make Evidence Hard to Skip shows how to make those evidence choices part of routine product work rather than a separate discovery phase.
Related books
Two books to
read next.
If you want to go further on this topic, these are two good places to start.
01
leadership
An Elegant Puzzle
by Will Larson
A human-centric guide to solving complex problems in engineering management, from sizing teams to handling technical debt to managing organizational growth.
02
leadership
The Five Dysfunctions of a Team
by Patrick Lencioni
A leadership fable about behaviours that damage teams and a practical model for rebuilding trust, conflict, commitment, accountability, and results.
Some outbound links are affiliate links and support independent bookstores.