Skip to content
Back to the journal

Essay

101

Design & Delivery

11 min read

101 / 136

Backlog Refinement: Govern Uncertainty Before Commitment

Refine product backlog items by exposing the uncertainty that matters, choosing the next evidence, and preserving what the team still does not know.

Updated July 13, 2026

Topics Team collaboration Delivery UX design

Share this essay

A backlog item can be impeccably written and still be a dangerous commitment.

The user story names an actor. The acceptance criteria are testable. Design has attached the happy path. Engineering has supplied an estimate.

Nobody has established whether the team may store the data the feature requires, whether support can reverse its effect, or whether the proposed behaviour addresses the problem that research uncovered.

The item looks ready because its visible fields are complete. The decision is weak because its consequential uncertainty is hidden.

Backlog refinement should make candidate work selectable under uncertainty. Its job is not to remove every unknown.

It should expose the unknowns that could change selection, scope, sequence, or the responsible way to build—and decide what to do about each one before commitment becomes expensive.

“Ready” is a claim, not a sticker

The 2020 Scrum Guide describes Product Backlog refinement as an ongoing activity that breaks work into smaller, more precise items.

It adds details such as description, order, and size. Items that the team can complete within one Sprint are considered ready for selection.

It does not prescribe a Definition of Ready, a refinement meeting, a ticket template, or a fixed number of Sprints to look ahead. Those are optional team practices.

That distinction matters. A checklist can help people remember recurring concerns, but it cannot decide whether the remaining uncertainty is acceptable for this item, this product, and this commitment.

Treat readiness as a bounded claim:

For this intended outcome, within these constraints, the people who may select and deliver the work understand the important evidence, decisions, dependencies, and unresolved uncertainty well enough to make a responsible commitment.

The claim is deliberately conditional. A production permission change and a disposable discovery prototype should not cross the same threshold. Nor should the threshold imply that learning will stop after selection.

Separate refinement from its neighbours

Several conversations meet at the backlog, which makes their purposes easy to blur.

  • Discovery asks whether a problem, opportunity, or proposed response deserves further investment.
  • Prioritisation decides which options should receive scarce capacity and what will be displaced.
  • Refinement determines whether a candidate option is understood well enough to be selected and what uncertainty must be resolved first.
  • Planning combines selectable work into a coherent goal within capacity and constraints.
  • Execution protects that goal while new information and operating conditions emerge.

Refinement should inform priority. It should not quietly reprioritise the portfolio because a detailed ticket feels more credible than a less-developed strategic bet.

How to Prioritise deals with that investment decision directly.

Refinement should reduce avoidable surprises before planning without pretending to predict implementation completely.

Product Planning: Commit Without False Certainty covers how candidate work becomes a revisable plan.

Open uncertainty in distinct lanes

“Any questions?” is a poor refinement method. It asks a group to search the entire problem space at once and rewards the first articulate answer.

Use several uncertainty lanes instead. Not every item needs work in every lane; the separation helps the team notice what a polished ticket has concealed.

Value uncertainty

What evidence connects this work to a user or business outcome? Which user, situation, and problem are in scope? What observation would weaken the case?

A feature request is evidence that somebody wants a feature. It is not yet evidence about the underlying job, the affected population, or the likely consequence of building it.

Interaction uncertainty

What must a user understand, choose, recover from, or hand over? Which states, roles, channels, and accessibility needs materially change the experience?

The happy path is rarely the whole interaction. Empty, delayed, denied, duplicated, interrupted, and partially completed states often carry more risk.

Technical and data uncertainty

Which systems, interfaces, identities, migrations, performance limits, or failure modes could change scope? What needs a spike, model, or production-like test rather than another discussion?

The purpose is not to dictate implementation. It is to find technical facts that could invalidate the proposed boundary or sequence.

Operational uncertainty

Who will observe, support, explain, reverse, and retire the change? What happens outside office hours? Which manual work or exception path does the new behaviour create?

An item may be easy to code and expensive to operate. Refinement should make that asymmetry visible before the implementation estimate becomes the apparent total cost.

Obligation uncertainty

Which privacy, security, safety, contractual, financial, regulatory, or policy conditions apply? Who has the authority to interpret and accept them?

Do not turn specialists into ceremonial approvers at the end. Bring the relevant expertise into the decision while the shape of the work can still change.

Give each unknown a disposition

Surfacing uncertainty without assigning a response produces an impressive list of worries and no safer decision.

For every consequential unknown, choose one disposition:

  • Resolve before selection: the answer could reverse the investment or materially change the work.
  • Bound before selection: agree a constraint, fallback, or excluded state that makes commitment responsible without a complete answer.
  • Carry into delivery: the team can learn safely during implementation; name the owner, observation point, and response.
  • Accept explicitly: the uncertainty remains, its exposure is understood, and a person with the right authority accepts it.
  • Stop or reshape: the evidence makes this version of the item irresponsible or incoherent.

“We will work it out in the Sprint” is not a disposition unless the team can say how, when, and what it will do if the answer is unfavourable.

The same applies to “blocked.” Name the missing decision or evidence, its owner, and the consequence of delay. A label does not create movement.

Match evidence to the decision deadline

Refinement performed too early converts guesses into detailed inventory. Performed too late, it converts uncertainty into delivery pressure.

Do not solve that problem with a universal rule such as “refine two Sprints ahead.” Use an evidence horizon.

Work backwards from the latest responsible selection date:

  1. Which unknowns could change whether the item should be selected?
  2. What is the cheapest credible way to resolve or bound each one?
  3. How long does that evidence or decision actually take?
  4. Which external lead times cannot be compressed?
  5. When must the first refinement action begin to preserve a real choice?

A data-retention interpretation may require early specialist input. A user-flow ambiguity may need a short prototype close to selection. A reversible copy decision can remain open during delivery.

Refine to preserve options, not to maximise ticket completeness.

Use conversations that fit the unknown

One weekly meeting with the whole team is only one possible coordination device. It is often the wrong place to resolve every issue.

Use the smallest useful group for the evidence or decision:

  • a product manager and researcher can test whether the problem boundary reflects current evidence;
  • a designer, engineer, and accessibility specialist can examine interaction states;
  • two engineers can investigate a risky integration and return with options;
  • product, operations, and support can define failure and recovery paths;
  • the full delivery group can challenge the resulting boundary before selection.

The full-team conversation remains valuable when perspectives must collide. It becomes wasteful when most participants watch two specialists discover facts they could have established beforehand.

Verwijs and Russo’s peer-reviewed mixed-methods study began with 13 case studies involving informants from more than 45 Scrum teams.

Its second phase tested the resulting model through a cross-sectional survey of 4,940 respondents across 1,978 Scrum teams.

In the qualitative observations, teams used different refinement arrangements. They repeatedly raised clarity, dependencies, and easier Sprint Planning as concerns.

That study does not prove one refinement format causes effectiveness. It supports the narrower point that refinement is a coordination problem teams must shape for their context, not a ceremony with one required design.

Preserve the decision, not the meeting transcript

The backlog item should help somebody recover the current decision without replaying every conversation.

Keep a concise refinement record:

Intended outcome and in-scope user situation
Evidence supporting the item
Chosen boundary and meaningful exclusions
Consequential unknowns and their dispositions
Dependencies, obligations, and decision owners
Acceptance evidence and operating safeguards
Remaining assumptions carried into delivery
Reason it is selectable now

Link to research, diagrams, decisions, and tests rather than pasting them all into the ticket. The ticket is a navigational record, not a substitute for a knowledge system.

GOV.UK’s user-story guidance emphasises the actor, need, and goal. It treats acceptance criteria as outcomes and recommends linking supporting evidence.

The guidance is specific to UK government digital services, not a universal ticket standard. Its useful lesson is that the artefact should keep the user goal and supporting evidence connected.

Quality requirements need deliberate attention too. Behutiye and colleagues interviewed 15 practitioners across four agile software-development cases and ran follow-up workshops in two of them.

The study found context-dependent documentation practices for quality requirements. The authors warn that such requirements can be under-specified; see the published paper record.

The authors say the findings are difficult to generalise beyond similar contexts. The study does not justify a universal non-functional checklist.

It does justify asking which quality conditions are consequential for the current item instead of assuming they will appear automatically.

A fictional refinement decision

Consider an explicitly fictional B2B product that lets finance teams approve supplier payments. A proposed item says: “Allow an approver to delegate approval while away.”

The item initially looks small. During refinement, the team separates its unknowns.

Value evidence suggests holiday coverage is a real source of delay, but it does not show that unrestricted delegation is the right response.

Interaction work reveals confusion when the original approver returns. Engineering finds that the current permission model cannot represent a temporary delegation.

Operations asks who can revoke access when the delegator is unreachable. Compliance needs an auditable distinction between delegated and original authority.

The team does not expand the ticket into a complete delegation platform. It chooses a boundary: one named substitute, a fixed end time, no onward delegation, visible provenance on every approval, and an administrator revocation path.

It resolves the permission-model question with a short technical investigation. It carries copy details into delivery with a design review point. It excludes recurring delegation and records that exclusion for future prioritisation.

No delivery result is claimed. The example shows refinement changing the decision boundary before the apparent simplicity of the request becomes a commitment.

Measure escaped uncertainty, not meeting activity

Ticket counts, refinement hours, and the percentage of items marked ready measure process volume. They do not show whether refinement protected decisions.

Inspect the unknowns that escaped their intended boundary:

  • work selected before a decision-changing dependency was known;
  • scope materially reshaped by evidence available before commitment;
  • obligation or operating work discovered only after implementation began;
  • items returned because the outcome or user situation was unclear;
  • assumptions carried into delivery without an owner or observation point;
  • detailed items discarded because refinement ran far ahead of selection.

Do not turn every change during delivery into a defect. Some learning is only available through building and use. The useful question is whether the team made a responsible choice about when and how to learn.

If problems emerge after selection, Sprint Execution explains how to protect the goal, expose ageing work, and make scope trades visible without lowering quality.

Run one refinement reset

Choose a candidate item that appears ready and ignore its status label.

Ask the people who may select and deliver it to identify one consequential unknown in each relevant lane. Give every unknown a disposition.

Then write a single sentence explaining why the item is selectable now—and under which conditions that claim would stop being true.

If the sentence depends on “we assume,” “someone will confirm,” or “we can handle it later,” the item may still be worth pursuing. It is not yet an honest commitment.

Sources

Related books

If you want to go further on this topic, these are two good places to start.

01

leadership

An Elegant Puzzle

by Will Larson

A human-centric guide to solving complex problems in engineering management, from sizing teams to handling technical debt to managing organizational growth.

Some outbound links are affiliate links and support independent bookstores.