Vanity Metrics: Choose Numbers That Can Change a Decision
Distinguish a useful product signal from a vanity metric by testing its decision, value mechanism, claim boundary, counter-signals, and incentives.
Piotr Ciechowicz
Product manager · developer
Updated July 13, 2026
On this page12 sections
- 01Vanity is a relationship, not a number
- 02Start with the choice that is still open
- 03Describe value before choosing the signal
- 04Keep the proxy one step below the claim
- 05Define the measurement boundary
- 06Use a small evidence set, not a solitary winner
- 07Test the metric before attaching a target
- 08Expect targeting to change behaviour
- 09A fictional example: invitation volume
- 10Run the review around claims, not colours
- 11Sources
- 12Read next
A dashboard can improve every week while the product gets worse.
Registrations rise because acquisition widened. Feature adoption rises because the old route disappeared. Daily activity rises because a broken workflow makes people repeat the same task.
None of those numbers is inherently useless. Each becomes dangerous when the team treats visible movement as evidence of value without checking the mechanism behind it.
A vanity metric is therefore not a fixed category such as page views, downloads, or active users. It is a measure used beyond the claim it can support.
The remedy is not a universal list of “metrics that matter.” It is a disciplined link between a product decision, a theory of value, observable evidence, and the conditions that would make the interpretation wrong.
Vanity is a relationship, not a number
Consider monthly active users.
For an annual tax-filing product, monthly activity may say little about value. For a product designed around a monthly close, it may be relevant. For an incident-response tool, more activity could mean adoption or more incidents.
The label does not settle the question.
A metric becomes vanity when at least one of these links is missing:
- the recurring decision it is meant to inform;
- the user or business outcome it represents;
- the mechanism connecting observed behaviour to that outcome;
- the population, unit, and time boundary of the claim;
- the counterevidence that could reverse the interpretation;
- the action available if the signal moves.
“Actionable” is not enough as a replacement label. A team can act on a misleading number. A metric deserves influence when the action and the inference are both defensible.
Start with the choice that is still open
Do not begin a metrics workshop by listing available events. Begin with a decision the team can still change.
Write it in a form with real alternatives:
At the next review, we will decide whether to continue, narrow, revise, investigate, or stop this commitment for this population.
If every plausible result leads to “continue,” the number is decoration. If nobody has authority to choose among the alternatives, it is reporting rather than decision support.
Then specify what the metric must distinguish. A team considering wider rollout may need evidence of repeated value and unacceptable failure, not a general engagement score.
A team diagnosing onboarding may need to know where eligible users fail to reach a meaningful state, not whether total registrations increased.
The data-informed decision guide covers how to preserve the wider evidence and uncertainty around that choice.
Describe value before choosing the signal
Product value is usually an unobservable construct. “Confidence,” “control,” “successful collaboration,” and “reduced risk” do not arrive as clean event fields.
Teams choose observable behaviour as evidence. That translation is where vanity enters.
Rodden, Hutchinson, and Fu introduced HEART as a framework for user-centred measures and paired it with a goals–signals–metrics process.
The process matters more here than the five HEART categories. It asks a team to state a product goal, identify behaviour or attitude that would signal progress, and only then define a metric.
The 2010 CHI note reports practical use inside Google. It is not a controlled validation that HEART improves decisions, and its examples concern large-scale web applications.
Its useful discipline is to keep the metric downstream from an explicit goal and signal.
For any proposed measure, write the chain:
Value claim
→ expected user or system change
→ observable signal
→ metric definition
→ decision the result can change
Now attack each arrow.
Could the behaviour occur without value? Could value occur without the behaviour? Could the event be generated by staff, retries, automation, coercive design, or a product defect?
Those questions do not invalidate a proxy. They define what else the team must observe before promoting the proxy into a verdict.
Keep the proxy one step below the claim
A saved workflow may indicate intent to use an automation. It does not prove that the automation ran, completed correctly, or saved meaningful effort.
A completed setup may be evidence of activation. It does not prove retention.
A renewal may indicate continued value. It may also reflect switching cost, contract timing, or an absent alternative.
Name the measure at the level it actually observes. “Workflows saved” is more honest than “customer value created” when saved workflows are the only evidence available.
Then label the relationship to the outcome as a hypothesis:
We expect organisations that successfully run a saved workflow to reach the intended operational result more often. We have not established that a save alone predicts that result.
This wording leaves room for validation. It also prevents a convenient leading signal from quietly becoming the strategy.
Define the measurement boundary
Even a well-chosen signal can mislead when its calculation is vague.
State:
- the unit: person, account, organisation, subscription, device, or transaction;
- eligibility and exclusions;
- the qualifying event or state;
- identity and deduplication rules;
- event time, reporting window, and maturity period;
- numerator and denominator;
- late data, backfills, and known coverage gaps.
A ratio is not automatically more meaningful than a count. Conversion can rise because the numerator improved or because fewer difficult cases entered the denominator.
Daily-to-monthly activity can change because daily use grew, monthly reach fell, or both moved.
Always show the components when their movement could change the decision.
The product metrics practice explains how to version these definitions, quality limits, owners, and migrations.
Choosing the right claim comes first; governing the calculation keeps that claim reconstructable.
Use a small evidence set, not a solitary winner
One metric rarely carries an important product decision safely. A useful set gives different evidence about the same commitment.
For a specific decision, consider four roles:
- Outcome evidence: the user or business state the commitment is meant to improve.
- Mechanism evidence: the behaviour expected to produce that state.
- Counter-signal: a cost, harm, displacement, or segment effect that could reverse the choice.
- Measurement health: evidence that the first three numbers are currently trustworthy.
These are roles, not a universal dashboard hierarchy. A metric may play different roles for different decisions.
Microsoft’s experimentation guidance uses a related distinction among overall evaluation criteria, local diagnostic measures, guardrails, and data-quality measures.
That guidance comes from Microsoft’s experimentation practice. It is not evidence that the taxonomy transfers unchanged to every product review.
Its valuable warning is that a local movement is easier to interpret when the team can also see the intended outcome, possible harm, and telemetry health.
Do not add metrics until every story has an opposite story. Add the smallest counter-signal capable of changing the interpretation.
Test the metric before attaching a target
A metric should earn authority before it receives a target, bonus, or executive promise.
Inspect historical cases where the product was known to improve, regress, or change in a meaningful way. Did the proposed measure move in the expected direction? Where did it miss or produce a false alarm?
Compare it with qualitative evidence and operational records. Examine important segments separately. Look for product changes that moved the metric without changing the claimed outcome.
For causal questions, use a credible comparison where one is ethical and feasible. Observational movement alone does not establish that the product change caused the outcome.
The SIGIR 2024 paper What Matters in a Measure? examines metric validity, cost, time, reliability, sensitivity, and interpretability in information-retrieval systems.
It is a research agenda grounded in search evaluation, not a validated product-metric scorecard.
Its broader contribution is the refusal to collapse metric quality into one property. A measure can be relevant but too slow, sensitive but unstable, or cheap but unable to distinguish meaningful alternatives.
Record those trade-offs before calling the metric “key.”
Expect targeting to change behaviour
Once a measure affects compensation, people may adapt to it. Similar pressure is plausible when status or resource allocation depends on the same number.
That does not require fraud. People may rationally optimise the visible representation while losing sight of the underlying strategy.
Choi, Hecht, and Tayler call this strategy surrogation: treating an imperfect performance measure as though it were the strategic construct itself.
Across two experiments, they found support for the prediction that incentive compensation increases surrogation, with a stronger effect for one measure than for multiple measures of a strategic construct.
The studies concern experimental management-accounting settings, not product teams in the field. They do not prove that a four-metric product set prevents gaming or restores good judgement.
They do justify a concrete review question:
If a capable team maximised this number while ignoring the intended outcome, what behaviour would we expect to see?
Test that failure path before the metric becomes a target. Preserve qualitative review, customer outcomes, and explicit exceptions where a single number would reward the wrong work.
Multiple metrics can still be gamed. The point is not to build a larger scorecard. It is to keep the strategic construct visible and make substitution easier to detect.
A fictional example: invitation volume
Consider a fictional collaboration product whose team is deciding whether to keep an onboarding change that encourages workspace invitations.
Invitation count rises quickly, so it is tempting to declare the change successful. The team instead writes the value claim: a new workspace can complete shared work with the right participants and permissions.
Invitations sent are mechanism evidence. They are not the outcome.
The team defines eligible new workspaces, removes staff and test accounts, separates unique recipients from repeated invitations, and waits long enough for a recipient to act.
Outcome evidence concerns completion of the shared task. Counter-signals cover unwanted invitations, permission corrections, and support contacts. Instrumentation health checks whether delivery and acceptance events join correctly.
No values or results are supplied in this example. The decision remains open.
If invitations rise while shared work and counter-signals do not support the mechanism, the team has learned that the visible action is a weak proxy in this context.
If the evidence aligns, it still supports only the defined population, period, and decision. It does not turn invitations into a universal measure of collaboration.
Run the review around claims, not colours
For each important metric movement, ask:
- What exactly changed, for which eligible population and mature period?
- Is the measurement path healthy enough for this use?
- Which value claim are we making from the movement?
- What mechanism connects the signal to that claim?
- Which competing explanations remain plausible?
- Which counter-signal or segment could reverse the decision?
- What choice changes now, and what evidence would reopen it?
A metric review should end with a decision, an investigation, or an explicit statement that the evidence is insufficient. “Keep watching” needs a named trigger; otherwise it is avoidance disguised as monitoring.
Useful product metrics are not the impressive numbers. They are the bounded measures that help a team make a better choice while keeping the distance between signal and value visible.
Sources
- Rodden, K., Hutchinson, H., & Fu, X. (2010). Measuring the User Experience on a Large Scale: User-Centered Metrics for Web Applications. Proceedings of CHI 2010. Introduces HEART and goals–signals–metrics through Google practice; not a controlled validation of decision quality.
- Microsoft Experimentation Platform. Patterns of Trustworthy Experimentation: During-Experiment Stage. Practitioner guidance on four metric types in controlled experiments; not a universal product-metrics standard.
- Thomas, P., Kazai, G., Craswell, N., & Spielman, S. (2024). What Matters in a Measure? A Perspective from Large-Scale Search Evaluation. Proceedings of SIGIR 2024. A perspective paper and research agenda for information-retrieval metrics, not a validated general product scorecard.
- Choi, J., Hecht, G. W., & Tayler, W. B. (2012). Lost in Translation: The Effects of Incentive Compensation on Strategy Surrogation. The Accounting Review, 87(4), 1135–1163. Two experiments in management-accounting settings; transfer to product organisations requires caution.
Read next
Related books
Two books to
read next.
If you want to go further on this topic, these are two good places to start.
01
leadership
An Elegant Puzzle
by Will Larson
A human-centric guide to solving complex problems in engineering management, from sizing teams to handling technical debt to managing organizational growth.
02
leadership
The Five Dysfunctions of a Team
by Patrick Lencioni
A leadership fable about behaviours that damage teams and a practical model for rebuilding trust, conflict, commitment, accountability, and results.
Some outbound links are affiliate links and support independent bookstores.