Skip to content
Back to the journal

Essay

080

Discovery & Validation

10 min read

080 / 136

Product Hypothesis Testing: Make Explanations Compete

Turn product hypotheses into competing explanations, discriminating evidence, explicit decision rules, and learning that survives an inconclusive result.

Updated July 13, 2026

Topics Validation User research Discovery

Share this essay

“If we add reminders, approval time will fall” looks like a hypothesis. It can still hide every difficult question.

Why are approvals slow? Which people and cases are affected? Will the intervention reach them? What else could produce the same observation? Which result would make the team choose differently?

A product hypothesis is useful when it makes explanations compete for a decision. A prediction-shaped sentence is not enough.

The work is to expose the causal story, derive observations that distinguish it from credible alternatives, and decide what each evidence pattern earns before the preferred story becomes persuasive.

Start with the decision the evidence can change

“Validate the hypothesis” has no stopping point. A team can always collect another quote, move a threshold, or explain away an inconvenient result.

Write the decision first:

  • the options that remain open;
  • the owner and decision date;
  • the cost and reversibility of each commitment;
  • the uncertainty most likely to reverse the current preference;
  • the evidence patterns that would support, weaken, or leave the choice unresolved;
  • the people, contexts, and product states in scope.

The decision determines how much evidence is proportionate. An editable copy change and an irreversible eligibility rule should not share a proof burden.

This keeps hypothesis testing narrower than concept validation. Concept validation decides whether a proposed response deserves another bounded commitment.

Hypothesis testing examines which account of a problem, mechanism, or expected change survives a discriminating check.

Replace one confident story with competing accounts

Teams often write one mechanism after choosing a feature:

People miss approvals because they are not reminded.

That account competes with others. Ownership may be unclear. The reviewer may lack authority. The request may need missing evidence. Existing notifications may be ignored because they contain no consequence or due date.

Write at least one credible rival and one account in which the proposed change does nothing. Do not add alternatives merely to fill a template.

For each account, state:

  • the actor, context, and observed phenomenon;
  • the proposed mechanism;
  • conditions required for that mechanism to operate;
  • observations expected if it is materially true;
  • observations awkward for it but plausible under a rival;
  • evidence that would make the account no longer decision-relevant.

In his 1964 essay “Strong Inference,” John Platt advocated alternative hypotheses, experiments with outcomes that exclude alternatives, and repeated refinement.

It is a methodological argument from science, not evidence that product teams using a template make better decisions. The transferable discipline is to ask which observation separates live accounts.

Separate the hypothesis stack

A feature bet usually contains several hypotheses. Combining them lets a failure in one layer masquerade as a verdict on all the others.

Use this editorial stack:

  1. Phenomenon: a bounded problem or behaviour exists in the stated population and context.
  2. Mechanism: a proposed factor materially contributes to that phenomenon.
  3. Intervention: changing a specific condition should alter the mechanism.
  4. Delivery: the product can deliver the intervention with sufficient exposure, fidelity, and reliability.
  5. Measurement: the chosen evidence can represent the expected change without redefining it.
  6. Decision: a stated evidence pattern warrants a particular commitment.

This is not a statistical model or universal standard. It is a way to locate what the evidence actually challenged.

If users do not see a reminder because notification permissions are absent, the intervention did not receive a fair test. If the metric counts an approval page load rather than a completed review, the measurement claim failed.

A flat result can leave the mechanism untouched. A positive result can still reflect novelty, selection, another simultaneous change, or an invalid outcome measure.

Derive a discriminating prediction

“We expect improvement” is compatible with almost any story. A discriminating prediction says how observations should differ if one account rather than another matters.

Suppose the reminder account predicts that delays cluster where reviewers have not re-entered the product after assignment. The unclear-ownership account predicts delays even after repeated visits, especially when several people can act.

Those predictions suggest different evidence and different interventions.

Use a prediction table:

ObservationReminder accountOwnership accountMeasurement or scope risk
Reviewer never returns after assignmentMore compatibleStill possibleLogin events may miss delegated review
Reviewer visits repeatedly but no one actsAwkwardMore compatibleRequest may be blocked for another reason
Explicit owner completes while peers observeWeak supportStronger supportAssignment may correlate with request quality

The table does not calculate truth. It forces the team to specify what would be surprising under each account and where the evidence could mislead.

Seek negative cases. A fast approval without a reminder, or a slow approval with one, may expose a missing condition.

Choose a method that can see the claim

There is no method called “a hypothesis test” that covers every product claim.

  • A phenomenon claim may need representative behavioural or operational evidence with a defensible denominator.
  • A mechanism claim may need episode reconstruction, observation, process tracing, or a causal design.
  • An interaction claim may need a prototype with realistic states, tasks, and participants.
  • A technical claim may need a spike, load test, failure injection, or security review.
  • An intervention effect may need random assignment or another credible comparison when causality drives the decision.
  • A policy or service claim may require end-to-end evidence beyond the interface.

Do not rank qualitative work below quantitative work. Rank a method against the claim it can and cannot distinguish.

GOV.UK guidance says user needs should be based on research rather than assumptions and focus on the problem rather than a proposed solution. Its scope is UK government services, not a general proof standard.

The useful boundary is that evidence about a need does not automatically validate one intervention.

The experiment-design guide covers randomisation, exposure, power, guardrails, and interpretation when a controlled comparison is appropriate.

Reserve “statistical hypothesis test” for the statistical claim

Product teams use “test” for interviews, prototypes, technical checks, and controlled experiments. That broad language is workable until it borrows certainty from statistics.

A statistical hypothesis test has a defined null hypothesis, alternative, model assumptions, test statistic, and rejection rule.

NIST describes it as a quantitative decision mechanism for asking whether evidence is sufficient to reject a conjecture about a process.

Failure to reject is not proof that the product has no effect. The test may lack power, the effect may be smaller than the detectable threshold, assumptions may fail, or exposure may be contaminated.

Likewise, a statistically detectable movement is not automatically a worthwhile product outcome. Effect size, uncertainty, guardrails, duration, affected populations, and opportunity cost still belong to the decision.

Do not attach a p-value to an analysis chosen after inspecting many outcomes and present it as if it confirmed a prediction made in advance.

Fix the interpretation rules before seeing the result

Write the evidence contract before collection or analysis:

  • primary claim and live alternatives;
  • population, unit, eligibility, and exclusions;
  • method, assignment or recruitment, and sample limits;
  • intervention version, exposure, and fidelity checks;
  • outcome, guardrails, measurement window, and maturity;
  • analysis choices and acceptable deviations;
  • stop conditions for harm, invalid data, or broken delivery;
  • actions attached to supporting, contradicting, and inconclusive patterns.

This does not prohibit learning from surprise. It stops surprise from being rewritten as a prediction.

The Center for Open Science describes preregistration as a way to distinguish confirmatory from exploratory analysis by recording a plan before data collection or analysis.

That is research-transparency practice, not a product-process requirement or guarantee of good reasoning. Product teams can borrow a lighter principle: preserve what was decided before the evidence and label later exploration honestly.

Read the failure path before the outcome metric

When an intervention shows no expected change, inspect the chain in order:

  1. Was the eligible population identified correctly?
  2. Was assignment, recruitment, or selection consistent with the design?
  3. Did people receive the intended intervention?
  4. Did the intervention behave with the required fidelity?
  5. Did the evidence window allow the outcome to mature?
  6. Did the measure represent the outcome for the affected population?
  7. Did guardrails reveal displacement or harm?
  8. Only then: what changed in the competing accounts?

Do not repair a broken test by declaring the product hypothesis false. Do not rescue an intact negative result by calling it a measurement issue without evidence.

An inconclusive result is a result about the current evidence design. The next action may be to improve the test, narrow the claim, choose a cheaper decision, or stop because the uncertainty is too expensive to resolve.

Preserve the hypothesis after the meeting

A hypothesis loses value when only the winning sentence survives. Keep a versioned record containing:

  • the decision and hypothesis stack;
  • competing accounts and discriminating predictions;
  • evidence contract and deviations;
  • observations separated from interpretation;
  • result, uncertainty, and scope;
  • decision, dissent, and rejected alternatives;
  • remaining hypotheses and reopening conditions;
  • expiry triggers such as a changed population, workflow, policy, or product state.

The experiment-tracking article explains how assignment, exposure, analysis lineage, deviations, and decisions should remain connected across repeated experiments.

Hypothesis records cover a broader set of evidence methods. They should link to, not duplicate, the authoritative research, telemetry, technical, or experiment record.

A fictional approval-delay investigation

Consider an explicitly fictional procurement product. The team observes that some supplier requests wait for review and proposes reminders.

It writes three accounts: reviewers forget, ownership is ambiguous, or requests arrive without evidence needed for a decision.

The accounts produce different predictions. Forgetting should be more common when no reviewer returns after assignment. Ambiguity should persist across visits where several people can act.

Missing evidence should produce requests for clarification, attachments added after assignment, or delays concentrated in particular request types.

The team checks instrumentation and reconstructs recent decision episodes before choosing an intervention. It defines what each evidence pattern would earn and where its records cannot distinguish delegation from inactivity.

If it later runs a controlled reminder test, exposure and delivery become separate hypotheses.

The primary outcome represents completed review within an independently defined eligible window, with delay displaced to clarification as a guardrail.

No conclusion is supplied here. The example shows how competing explanations can prevent one convenient feature from defining the problem.

A good hypothesis makes a decision easier to reverse

The standard is not whether a hypothesis sounds precise. It is whether another team can see the alternatives, inspect the evidence boundary, reproduce the interpretation, and understand why the decision followed.

Good hypothesis testing can support, weaken, refine, or leave an account unresolved. It can expose a broken intervention, an invalid measure, or a decision that does not justify further evidence cost.

Its product value comes from making the next commitment smaller, clearer, and easier to reverse when the world contradicts the story.

Sources

Related books

If you want to go further on this topic, these are two good places to start.

01

leadership

An Elegant Puzzle

by Will Larson

A human-centric guide to solving complex problems in engineering management, from sizing teams to handling technical debt to managing organizational growth.

Some outbound links are affiliate links and support independent bookstores.