Skip to content
Back to the journal

Essay

044

Discovery & Validation

13 min read

044 / 136

Continuous Research: Manage Evidence, Not an Interview Quota

Run continuous research as an evidence flow: prioritise open decisions, control research WIP, cover missing users, expire stale findings, and change work.

Updated July 13, 2026

Topics User research Validation Discovery

Share this essay

A product team can interview two customers every week and still learn too late. The calendar stays full, summaries accumulate, and consequential decisions are made from memory, urgency, or the last persuasive conversation.

The problem is not insufficient consistency. It is that a cadence has been mistaken for an operating system.

Continuous research is better understood as a managed flow of evidence serving a portfolio of open product decisions. That changes the unit of management.

Instead of asking, “Did we speak to users this week?”, ask which decisions need evidence, who remains unseen, what the team can analyse responsibly, and whether existing findings still fit the decision.

A healthy cadence may emerge from that system. It is not the purpose of the system.

First, separate the practices people bundle together

“Continuous discovery” and “continuous research” are often used loosely enough to hide important differences.

Customer discovery examines who has a problem, how it appears in context, and whether it is important enough to address. It is useful when the problem, audience, or value proposition remains uncertain.

Discovery techniques are methods selected for a particular uncertainty. Interviews, observations, prototype tests, log reviews, surveys, and experiments answer different kinds of questions.

The right choice depends on the claim the team needs to make. A broader guide to product discovery techniques covers that selection problem.

Unsolicited feedback is an incoming signal: a support ticket, sales objection, review, cancellation note, or feature request. It can create a research question.

On its own, feedback does not establish who else has the problem, why it occurs, or which response would help.

Interview craft concerns one method: whom to recruit, how to avoid leading questions, how to probe past behaviour, and how to interpret what was said.

Interview quality matters. It does not decide whether an interview was the right method or whether that evidence deserves priority.

Continuous research is the operating layer across these activities. It governs priority, method, sample, work in progress, evidence provenance, and the route to a decision.

If these practices are collapsed into “talk to customers regularly”, a team may become very good at producing conversations while remaining poor at resolving uncertainty.

Run a research demand queue, not a request inbox

Most research requests arrive as methods: “We need five interviews”, “Can we test this screen?”, or “Let’s send a survey.” The method is being chosen before the research demand is legible.

A useful demand item should name:

  • the open decision and the person accountable for making it;
  • the deadline or commitment that gives the decision a time boundary;
  • the uncertain belief that could change the decision;
  • the consequence of being wrong or of waiting;
  • the evidence already available, including its source, scope, age, and weaknesses;
  • the people, situations, and product conditions relevant to the question;
  • the claim the proposed research could and could not support.

This is stricter than accepting a request for “some customer insight”. It also makes prioritisation possible.

A reversible copy decision due next month should not displace an irreversible platform decision due next week merely because its sponsor asked first.

The GOV.UK planning guidance applies similar logic to government service teams: prioritise questions, then connect methods and user groups to them.

It also warns against producing more research than the team can respond to. This is context-specific practice guidance, not experimental proof or a universal cadence.

Review the queue when decisions, risks, or access conditions change. It is not a backlog to exhaust.

Items can be combined, reframed, answered with existing evidence, routed elsewhere, or closed because the decision is no longer open.

Protect several lanes of learning

If every study serves the next delivery decision, a team may miss changes in the market or in product use. If everything is open-ended exploration, urgent decisions wait for nonessential learning.

A practical portfolio can contain four lanes:

  1. Decision-bound research addresses an explicit choice with an owner and deadline.
  2. Product-health research investigates recurring friction, adoption changes, service failures, or contradictory signals in a live product.
  3. Coverage and assumption debt revisits important populations, contexts, and beliefs that have been persistently under-examined.
  4. Strategic exploration looks beyond the current roadmap for shifts in behaviour, constraints, alternatives, and emerging problems.

There is no universal allocation across these lanes. The balance depends on product maturity, decision load, risk, and access to evidence.

Make the trade-off visible. Otherwise, the loudest delivery request consumes every available research slot.

Some issues should not wait in the ordinary queue. A credible accessibility, safety, privacy, security, or rights concern may require immediate action and specialist escalation.

Research can clarify what happened. It should not postpone an obligation the team already understands.

Set WIP by analysis capacity

Research WIP includes recruitment, consent, fieldwork, data preparation, analysis, specialist review, decision discussion, and responsible storage or disposal.

Collection is the most visible part, so it is easy to schedule beyond the team’s ability to interpret. The result is evidence inventory: recordings and dashboards that exist but have not been examined closely enough to inform a choice.

A qualitative study by Tkalich and colleagues examined user-feedback practice through 21 interviews with 19 practitioners across 13 software-product cases.

The cases were linked to companies with offices in Norway. Recruitment used industrial contacts and referrals, with one to three informants per case. The authors did not seek population generalisation.

In this sample, practitioners reported constraints involving resources, analysis, metrics, access to suitable users, and links to planning. These accounts suggest failure modes, not their prevalence, causes, or an optimal cadence.

Useful WIP policies are operational rather than numerical:

  • no decision owner, no decision-bound study;
  • no analysis capacity, no new collection;
  • no longer an open decision, close or reframe the work;
  • a method that cannot support the required claim is rerouted;
  • blocked work receives an explicit review rather than ageing invisibly.

A team might still choose a numerical WIP limit, but that number should follow observed capacity. It should not be borrowed from another organisation’s ritual.

Track coverage debt before counting participants

More interviews do not necessarily broaden evidence. Ten conversations with available administrators can deepen that context while leaving operators, churned customers, buyers, non-users, and assistive-technology users unseen.

Call that persistent gap coverage debt: repeated dependence on accessible sources while groups or conditions relevant to a live decision remain missing.

A lightweight coverage register can record:

  • the decision or assumption being examined;
  • the roles, contexts, behaviours, product states, and access needs that matter;
  • which of them have and have not been observed;
  • when the evidence was gathered and against which product version;
  • how participants or data were sourced;
  • what the gaps prevent the team from claiming.

Coverage is not a demographic checklist detached from the question. A permissions decision may turn on role and authority. Onboarding may turn on prior knowledge and account state.

In an international workflow, language, regulation, channel, or operating environment may change what is possible.

The point is not to declare a qualitative sample “representative”. It is to expose the evidence boundary.

The Tkalich paper discloses non-probability recruitment, uneven informants per case, and a Norway-linked context, then limits its claims. Product teams owe their own evidence the same honesty.

Give evidence a lifecycle and an expiry trigger

A research repository is not a library of timeless facts. A finding is an interpretation produced under particular conditions.

An evidence item should retain enough provenance to answer basic questions later:

  • What question was being investigated, with which method?
  • Who or what was observed, when, and in what product and market context?
  • Who conducted the work, and where are the permitted source records?
  • What consent, privacy, access, and retention constraints apply?
  • What claim did the evidence support, and what did it not establish?
  • What conflicting or subsequent evidence exists?

That item can then move through a visible lifecycle: captured, interpreted, attached to a decision, active for a defined scope, and eventually superseded, expired, or archived.

Expiry should not mean “false after 90 days”. There is no universal shelf life. Define review triggers instead.

Reassess evidence when the product, population, buyer, channel, regulation, instrumentation, or classification changes, when credible contradictions appear, or when a decision asks a different question.

Expired evidence can still provide history and help shape a new question. It should not silently carry the authority of current evidence.

Give the team exposure without pretending everyone is a researcher

Research should not become a report-delivery service. Colleagues can observe sessions, inspect evidence, and join facilitated analysis. But exposure does not confer research competence.

The roles remain distinct:

  • the researcher is accountable for methodological quality, recruitment logic, consent, facilitation, analysis, and appropriate claims;
  • the product manager or other decision owner is accountable for making the demand explicit and recording what changes;
  • designers, engineers, support, sales, and domain specialists can contribute questions, observe where appropriate, inspect artefacts, and challenge interpretations;
  • legal, privacy, accessibility, safety, and security specialists own boundaries that cannot be settled through a team research ritual.

GOV.UK’s research introduction recommends small research batches during service development and team involvement in observation and analysis.

Its staffing, cadence, and exposure figures belong to that context. They are not universal performance targets.

GitLab documents one company practice.

Its Continuous Interviews handbook describes PMs using a largely stable script and inviting colleagues to attend.

Consented recordings, tags, and summaries connect to a shared repository. The page says these interviews are not organised around a specific hypothesis or research request.

That can serve an exploration and exposure lane. It is not a complete research system, and the handbook does not evaluate its effectiveness.

Attach evidence to a decision before collection begins

Before starting decision-bound research, write down:

  • the decision, owner, and date by which it must be made;
  • the belief or uncertainty under examination;
  • why the chosen method and sample can inform it;
  • what would support, weaken, or leave the belief unresolved;
  • what action is available in each case.

This links a research question to the product discovery questions that shape a decision. It prevents a post-hoc search for evidence that justifies favoured work.

After analysis, record whether the team proceeded, narrowed the commitment, stopped, sought different evidence, or reopened the decision. Also record what remained unknown. “We found an interesting insight” is not a decision outcome.

Not all research has an immediate product decision. Strategic exploration can create better questions or revise a market model.

It should still have a named audience, a review point, and a clear account of what the evidence can change. Otherwise, “exploration” becomes exempt from prioritisation.

Measure flow and use, not activity theatre

Interview count, repository volume, and research hours are easy to report. They say little about whether evidence arrived in time, covered the right conditions, or altered a commitment.

A research operating review can instead examine:

  • the age of research demand relative to decision deadlines;
  • WIP by state, including time blocked or awaiting analysis;
  • coverage gaps relevant to active decisions;
  • active decisions relying on evidence past a review trigger;
  • decisions that proceeded, narrowed, stopped, or reopened after evidence was considered;
  • unresolved questions with no accountable owner.

These measures are diagnostic, not quotas. A high number of stopped initiatives could indicate valuable early learning, indiscriminate stopping, or a badly formed portfolio. The record prompts examination; it does not supply the verdict.

Research is not the only stream of customer evidence. Support, sales, operations, market signals, and product behaviour reveal different parts of reality.

A broader customer-centric operating model connects them without pretending they are interchangeable.

Know when to stop or change the research

Continuous research does not mean endless research. Close or pause work when:

  • the decision is no longer open;
  • the evidence is sufficient for the size and reversibility of the current commitment;
  • the selected method cannot resolve the uncertainty;
  • no owner or analysis capacity exists;
  • access, consent, safety, or privacy constraints prevent responsible work.

Change the question, method, sample, or lane when evidence reveals a different population, mechanism, risk, deadline, or next commitment.

Stopping a study is not a declaration of certainty. It means more evidence from that work is not currently worth its cost or cannot support the required claim.

A fictional triage example

Consider a fictional workflow-software team with four demands: removing manual approval, a screen-reader failure in that flow, repeated bulk-export requests, and a longer-term question about cross-region handovers.

The accessibility report bypasses the ordinary queue for specialist investigation. The approval decision enters the decision-bound lane with an owner, deadline, relevant roles, and current evidence.

Export requests remain feedback intake until the team clarifies their contexts. The cross-region question keeps a bounded place in exploration rather than being erased by delivery work.

The approval evidence comes only from administrators and predates a permissions change. It is not automatically wrong, but it has crossed a review trigger and exposes missing approvers.

Because analysis capacity is already occupied, the team does not start every proposed study at once.

This example is fictional. It illustrates the mechanics of triage and makes no claim about the result such a team would achieve.

The operating review that keeps the system honest

A useful recurring review is short enough to survive but demanding enough to expose neglected work:

  1. Which product decisions are open, and when do they become costly to delay?
  2. Which uncertainties could genuinely change those decisions?
  3. What evidence is active, and which findings have crossed a review trigger?
  4. Which relevant users or conditions remain absent?
  5. Where is research WIP blocked, especially before analysis or decision use?
  6. What should start, stop, change lanes, or leave the queue?
  7. What decision changed, and what remains unresolved?

The test of continuous research is not whether interviews recur. It is whether evidence is replenished, challenged, scoped, attached to decisions, retired when stale, and used while a decision can still change.

Sources

Product Discovery Techniques and Strategies helps match the method to the uncertainty and the claim.

9 Product Discovery Questions That Force a Decision defines the decision frame before evidence is collected.

Practical Customer-Centricity for Product Teams connects research with feedback, service, behaviour, and commercial evidence.

Related books

If you want to go further on this topic, these are two good places to start.

01

leadership

An Elegant Puzzle

by Will Larson

A human-centric guide to solving complex problems in engineering management, from sizing teams to handling technical debt to managing organizational growth.

Some outbound links are affiliate links and support independent bookstores.