Agile Estimation: Name the Decision Before You Estimate
Choose decomposition, reference-class forecasting, or a bounded investigation by the decision an estimate must support and the uncertainty it can reduce.
Piotr Ciechowicz
Product manager · developer
Updated July 13, 2026
On this page11 sections
- 01Start with the decision, not the unit
- 02Keep estimation between its neighbours
- 03Classify the uncertainty before selecting a method
- 04Decompose until the decision becomes stable
- 05Forecast from comparable completed work
- 06Buy information when estimation would be theatre
- 07Refuse false conversion
- 08A fictional estimate that changes the work
- 09Preserve the estimate as a revisable claim
- 10Sources
- 11Read next
A stakeholder asks when an initiative will be done. The team answers with story points. Finance hears a date, engineering hears relative effort, and product hears a reason to keep the initiative small.
The number is not merely inaccurate. It is answering several incompatible questions.
Estimation becomes useful when it is treated as a decision instrument. First name the choice that needs support. Then select a method that fits the work, the available evidence, and the cost of being wrong.
Sometimes the right instrument is decomposition. Sometimes it is a reference class built from completed work. Sometimes the only honest estimate is a bounded step that buys missing information.
No technique deserves to predict all three situations.
Start with the decision, not the unit
“We need an estimate” is an incomplete request. Ask what somebody will do differently after receiving it.
Common decisions include:
- whether a candidate is understood well enough to compare with alternatives;
- whether a team can protect a near-term goal while accepting this work;
- whether a release window or external promise remains credible;
- whether an unknown deserves investigation before a larger commitment;
- whether to stop, split, sequence, or narrow the proposed work.
Write a short estimation brief before choosing points, days, sizes, or simulations:
Decision and decision owner
Object being estimated
Boundary of “finished”
Latest useful decision date
Evidence currently available
Consequence of under- and over-estimation
Uncertainty that could reverse the decision
Trigger for revising the estimate
The cost of error should shape the answer. A reversible sequencing choice may tolerate a broad comparison. A contractual date may require a probability range, explicit exclusions, and an escalation path.
Precision should follow the decision burden. It should not follow the number of people in the meeting.
Keep estimation between its neighbours
Estimation is easily asked to compensate for missing product and delivery decisions.
Prioritisation decides which options deserve investment and what must give way. Effort evidence informs that choice, but effort alone cannot determine value or strategic fit.
Product planning reconciles selected bets, obligations, capacity, and dependencies. An estimate is one input; it is not the commitment system.
Backlog refinement exposes unknowns and gives them a disposition before selection. Estimation examines how those unknowns affect effort or timing.
Sprint execution protects a goal while scope, flow, and operating conditions change. It should not be judged by whether every item matches its original estimate.
Portfolio roadmapping sequences bets and dependencies across a wider system. Adding team estimates does not resolve strategic conflict or shared capacity.
Discovery and experiments test product claims. A technical investigation used for estimation should not quietly become evidence that users need the product.
Classify the uncertainty before selecting a method
Three different conditions often receive the same estimate request.
| Condition | Useful instrument | Honest output | Typical misuse |
|---|---|---|---|
| The work is broadly understood but too coarse | Decomposition and structured team judgement | Relative size, effort range, assumptions and risky parts | Treating consensus as proof |
| Similar work moves through a reasonably stable system | Reference class or flow forecast | Distribution, probability, class definition and observation window | Turning a percentile into a guarantee |
| A feasibility or design fact is genuinely unknown | Bounded investigation | Evidence, remaining unknowns and a revised decision | Giving the investigation a disguised delivery date |
These instruments can be combined, but their outputs should remain separate. A technical spike can change the work boundary; a reference class can then forecast the newly bounded work.
Decompose until the decision becomes stable
Decomposition is valuable when a large label hides distinct work: migration, permissions, failure handling, observability, rollout, support, or retirement of an old path.
Start with the finish condition. “Code complete,” “available behind a flag,” and “usable by the intended population with a recovery path” describe different objects.
Then split the work along boundaries that can change the estimate or the decision:
- user and operator states;
- system interfaces and data movement;
- obligations and quality conditions;
- external decisions and dependencies;
- rollout, reversal, support, and removal.
Do not decompose merely to produce more tickets. Stop when another split would not change selection, sequencing, ownership, risk treatment, or the forecast.
If several practitioners estimate together, collect initial judgements independently. Discuss the reasoning behind the widest differences before seeking a shared range.
Ask what each person included, which reference they used, and which failure state they assumed away. The useful product of the discussion is a changed model of the work, not ceremonial unanimity.
Mahnič and Hovelja had 13 four-person student teams and a four-person expert group estimate 28 user stories. The analysis used estimates from 10 student teams.
The students developed the same system over three Sprints.
For students, group estimates became slightly more optimistic and less accurate than averaged first-round estimates. The experts’ group estimates were closer to actual effort and tended to improve on their individual combination.
This does not establish that Planning Poker works for experienced product teams. It shows that group discussion can help or harm depending on expertise and context, in one academic project with a small expert group.
Use simultaneous estimates to expose different models. Do not cite the ritual itself as evidence of accuracy.
Forecast from comparable completed work
When the decision concerns completion time or the amount of work likely to finish, start outside the current item’s story.
Define a reference class before looking for a convenient precedent. Useful boundaries may include work type, finish condition, service class, team and workflow, dependency pattern, or technical environment.
Then inspect completed work:
- Use the actual start and finish boundaries that match the new forecast.
- Preserve the distribution instead of reporting only an average.
- Document exclusions before seeing whether they improve the answer.
- State the sample period and any material change to team or workflow.
- Express the result as a probability or range tied to the decision.
- Update the forecast when work, capacity, or the system changes.
For a batch, a simulation can sample observed throughput or cycle times to produce a range of completion outcomes. The model does not remove uncertainty; it makes its assumptions inspectable.
The May 2025 Kanban Guide defines a service level expectation as elapsed time plus a probability and says it should be based on historical cycle time.
It also requires four flow metrics: WIP, throughput, work item age, and cycle time.
That is official Kanban guidance, not evidence that any team’s history is stable or comparable. A forecast is weak when work classes are mixed, finish states drift, or a bottleneck has materially changed.
Servranckx, Vanhoucke, and Aouam interviewed 76 project managers about project similarity, then evaluated reference-class forecasting with 52 completed real-world projects and cross-validation.
They found that using more similarity properties could improve accuracy, but also made reference classes smaller and less reliable. They explicitly warn that improvement is not guaranteed for every project.
Their sample spans project management rather than agile software teams. The transferable lesson is the trade-off: a broad class may be irrelevant, while a perfect match may contain too little evidence.
Buy information when estimation would be theatre
Some unknowns are not difficult quantities. They are missing facts.
An unfamiliar data format, undocumented vendor behaviour, uncertain migration path, or untested performance boundary cannot be made predictable by adding more points.
Use a bounded investigation when its result can change a named decision. Specify:
- the uncertainty being tested;
- the decision it will inform;
- the smallest credible environment or artefact;
- the evidence to capture;
- the stopping boundary;
- the owner who will interpret the result;
- possible dispositions after the investigation.
Possible dispositions are proceed, reshape, choose another path, gather different evidence, or stop. “Continue researching” needs a new question and authorisation.
The US Government Accountability Office’s Agile Assessment Guide describes a spike as research used when a story is hard to estimate because of a design or technical challenge.
Federal auditors are its primary audience; organisations and programmes can also use it. It is guidance, not an experiment showing that spikes improve forecasts.
Its useful boundary is that the spike exists to understand the work. It should not be sized as though its output and the eventual implementation are the same deliverable.
Refuse false conversion
A relative size is not a duration until it is combined with evidence about a particular team’s capability, capacity, workflow, finish condition, and uncertainty.
The 2020 Scrum Guide makes Developers responsible for sizing and says attributes vary by domain. It names past performance, upcoming capacity, and the Definition of Done as inputs to a Sprint forecast.
Scrum does not prescribe story points, Planning Poker, a Fibonacci scale, or conversion from points to hours. Those are optional practices.
Do not compare points between teams, use velocity as an individual target, or improve a forecast by pressuring the size downward. Such incentives change the measure before they change the work.
Keep three statements visibly separate:
- Size judgement: how this work compares with a defined reference under stated assumptions.
- Flow forecast: what a historical distribution suggests about completion.
- Commitment: what an authorised owner accepts after considering value, capacity, risk, and consequence.
When a range is unacceptable, the response is a decision: narrow scope, change sequence, add evidence, alter capacity, renegotiate the promise, or decline it. The arithmetic cannot choose.
A fictional estimate that changes the work
Consider a fictional B2B billing product. A customer renewal depends on exporting invoices from both the current ledger and an older acquired system.
The request arrives as “estimate the export.” The decision is sharper: can the company responsibly offer a delivery window before the renewal negotiation closes?
The team defines finished as an auditable export for both ledgers, including permission checks, reconciliation, failed-record reporting, and an operator recovery path.
Decomposition shows that the current ledger resembles completed export work. The acquired system does not: its adjustment records lack reliable documentation.
The team does not hide that unknown inside a larger size. It proposes a bounded investigation using a controlled data sample to test whether adjustments can be reconstructed and reconciled.
After that evidence, the team can reshape the export boundary. It can use comparable completed items for the known path and a flow forecast for the resulting batch, with the class and probability stated.
The commercial owner still decides what to promise. The estimate has supported the decision by separating known work, historical evidence, and unresolved feasibility.
This example is entirely fictional and claims no delivery result.
Preserve the estimate as a revisable claim
Keep the decision brief, work boundary, method, reference class, assumptions, range, confidence language, exclusions, and revision trigger together.
When reality differs, review the mechanism before scoring the estimator.
Was the finish condition changed? Did an excluded dependency enter? Was the reference class poor? Did WIP or capacity shift? Did the investigation answer the wrong uncertainty?
Calibration means improving the relationship between evidence and decisions. It does not mean training people to produce numbers that survive a status meeting.
An estimate earns its cost when it changes a choice, reveals a condition, or buys information. If nobody can name the decision, do not estimate yet.
Sources
- The Scrum Guide, November 2020 — official framework guidance. It assigns sizing to Developers and identifies inputs to Sprint forecasts; it does not prescribe story points or validate an estimation technique.
- The Kanban Guide, May 2025 — current official Kanban guidance on flow metrics and probabilistic service level expectations. It does not establish that a local dataset is stable or comparable.
- Mahnič and Hovelja, “On using planning poker for estimating user stories” — study in which 13 four-person student teams estimated 28 stories over three Sprints; the analysis used 10 teams, alongside four experts. Its academic setting and small expert sample limit transfer to professional teams.
- Servranckx, Vanhoucke, and Aouam, “Practical application of reference class forecasting for cost and time estimations” — 76 interviews and cross-validated analysis of 52 real projects. It concerns project forecasting broadly, not agile software delivery.
- GAO Agile Assessment Guide, reissued December 2023 — US federal assessment guidance that describes a spike as research for hard-to-estimate design or technical questions. It is practitioner guidance, not causal evidence.
Read next
Related books
Two books to
read next.
If you want to go further on this topic, these are two good places to start.
01
product
Continuous Discovery Habits
by Teresa Torres
A practical guide to discovering products that create customer value and business value, with frameworks for integrating customer research into weekly rhythms.
02
product
The Lean Startup
by Eric Ries
How today's entrepreneurs use continuous innovation to create radically successful businesses, introducing Build-Measure-Learn and validated learning.
Some outbound links are affiliate links and support independent bookstores.