Skip to content
Back to the journal

Essay

082

Design & Delivery

11 min read

082 / 136

Rapid Prototyping: Build the Evidence, Not the Product

Choose prototype fidelity from the decision and risk, test the smallest credible model, and end with evidence, limits, and an explicit next move.

Updated July 13, 2026

Topics Delivery UX design Engineering

Share this essay

A clickable prototype gets an enthusiastic response in a review. One person hears proof of demand. Another hears approval of the navigation. Engineering hears a commitment to build what was shown.

The prototype did not fail. The team failed to say what it could distinguish.

A useful prototype is the cheapest credible model that separates competing decisions. It does not need to resemble the final product in every way. It needs enough reality in the places that could change the decision.

Fidelity therefore follows the question and the risk. It is not a ladder that every idea climbs from sketch to polished software.

Name the decision before the artifact

Begin with alternatives, not a tool. Write the decision that will follow the work:

  • proceed with one interaction model or another;
  • investigate a technical mechanism or reject it;
  • expose the idea to a narrower audience or stop;
  • automate a service step or keep a person in the loop.

Then state the uncertainty preventing that decision. “We need feedback” is not an uncertainty. “We do not know whether operators can detect and correct a proposed classification before it affects a customer” is.

The prototype claim should connect the representation to observable evidence:

If people with the relevant experience attempt this task in a credible context, this prototype should reveal whether they can notice, understand, and recover from the proposed system behaviour.

Add the exclusion beside the claim. The same prototype may say nothing about demand, model accuracy, production latency, operational cost, or accessibility outside the tested path.

This keeps rapid prototyping distinct from concept validation. Concept validation asks whether a proposed response deserves another commitment.

A prototype represents selected qualities so a team can examine them.

Treat fidelity as a profile

“Low fidelity” and “high fidelity” conceal more than they explain. A prototype can be visually crude and technically real, or visually exact while every result is supplied by a person behind the screen.

Houde and Hill proposed looking at a prototype through its role, look and feel, and implementation.

They also warned that prototypes are not self-explanatory. Audiences need to know which parts correspond to the intended artifact.

Lim, Stolterman, and Tenenberg later described prototypes as filters and manifestations of design ideas.

Their anatomy is descriptive. The paper reinterprets two earlier case studies; it does not establish a universal process standard.

For product work, write a fidelity profile across the dimensions that matter:

  • Scope: which journey, branch, state, or handoff exists, and what is absent?
  • Visual: which hierarchy, density, content, brand, and accessibility cues are credible?
  • Interaction: which inputs, transitions, errors, interruptions, and recovery paths behave realistically?
  • Data: is the shape, variation, messiness, volume, and sensitivity of the data representative?
  • Technical: which integrations, algorithms, latency, persistence, security, and failure modes are real?
  • Environment: does the device, network, location, noise, time pressure, and surrounding workflow matter?
  • People and service: which human decisions, support steps, approvals, and manual operations are represented?

Do not average these dimensions into one score. Raise a dimension only when its current simplification could produce the wrong decision.

Draw the critical realism boundary

Every test has a point beyond which a shortcut stops being harmless. That point is the critical realism boundary.

If the question concerns comprehension, realistic terminology and information may matter more than polished colour. If it concerns a slow connection, a smooth local animation is not credible evidence.

If trust depends on who approves an exception, replacing that person with a generic success screen removes the mechanism under study. If privacy affects behaviour, using implausibly clean dummy data can make the task easier than reality.

Write two lists before building:

  • Must be credible: qualities that could reverse the decision if represented badly.
  • Deliberately simulated: qualities excluded from the claim, with the likely direction of distortion.

The second list matters. A prototype that hides its substitutions invites people to promote a narrow observation into a broad conclusion.

Choose the cheapest credible evidence

The fastest prototype is not always software. It is the least expensive representation that can still expose the mechanism behind the uncertainty.

  • Fake door: reveals whether an eligible audience attempts to enter a clearly stated proposition at a real decision point. It cannot show delivered value, retention, or whether the service works.
  • Paper prototype: exposes terminology, sequence, grouping, and visible choices. It cannot support claims about timing, responsive behaviour, assistive technology, or system feedback unless those are represented credibly.
  • Clickable prototype: examines navigation, hierarchy, and interaction cues. It usually hides data variation, latency, concurrency, permissions, and failure recovery.
  • Wizard of Oz: lets a hidden person simulate system behaviour. It can examine the experience around a capability, but not whether the automation is feasible, accurate, scalable, or affordable.
  • Concierge prototype: openly delivers a proposed service through people. It can reveal workflow and value boundaries without pretending the service is automated. It does not establish scalable unit economics.
  • Functional prototype: uses working code where runtime behaviour matters. Its evidence is limited by the realism of its data, integrations, environment, load, security, and operational support.
  • Technical spike: isolates an implementation uncertainty. It can show that a mechanism works under stated conditions. It cannot show that people need, understand, or value the resulting product.

These are working forms, not maturity stages. A paper flow can follow a technical spike. A concierge service can remain the honest product choice if automation adds risk without useful value.

GOV.UK guidance ranges from sketches to code prototypes and recommends selecting the form that meets the current need. Its rules address government services; the underlying warning about copying prototype code into production travels well.

Put the task in its real context

Recruitment is part of prototype validity. A fluent colleague can make an unfamiliar workflow look obvious. A domain expert can complete a task while silently compensating for a broken design.

Recruit people whose experience, responsibility, access needs, and constraints match the decision. There is no universal participant count.

The coverage needed depends on variation in the audience, the consequence of error, and what the next commitment costs.

GOV.UK advises recruiting actual or likely users and setting criteria from the research questions. That is public-service guidance, not a statistical rule for every product study.

Context deserves the same care. Test on the device, data shape, channel, and location that could alter behaviour. Include the handoffs and interruptions that the claim depends on.

For a moderated test, give participants a believable goal without naming the interface action that completes it. GOV.UK’s usability guidance similarly recommends relevant tasks that do not reveal the answer.

Separate observation from interpretation

A prototype session produces traces, not a verdict. Preserve the path from what happened to what the team decides.

Use a simple evidence record:

  • Observation: what the participant did or said, with the task state and prompt recorded.
  • Interpretation: the proposed reason, written as an inference rather than a fact.
  • Alternative explanation: another cause compatible with the same observation.
  • Decision relevance: which alternative becomes more or less credible, within the prototype boundary.

“The participant returned to the previous screen after reading the permission warning” is an observation. “The warning destroys trust” is an interpretation.

The wording, scenario, prior experience, or missing detail may also explain the behaviour.

Use neutral prompts. Ask what the person expects next, what they believe changed, and how they would recover. Avoid selling the idea, correcting the task too early, or treating preference as evidence of successful use.

If the team wants a causal estimate, it needs an experiment with an assignment and measurement design suited to that claim. A handful of moderated sessions cannot establish production impact.

Decide the conditions in advance

Agree what evidence would change the decision before the prototype becomes persuasive. This is not a universal scoring framework; it is a defence against moving the standard after seeing attractive results.

  • Stop the session if consent becomes uncertain, the task creates unexpected harm, sensitive data escapes the agreed boundary, or the prototype no longer represents the claim.
  • Advance when the necessary behaviour is credible in the intended context and the unresolved risks fit the next bounded commitment.
  • Revise when the representation, task, or recruitment prevented a fair test of the claim.
  • Abandon when evidence contradicts a necessary premise, the risk cannot be contained, or another option now dominates.

An inconclusive result is not an instruction to add polish. It may mean the claim is too broad, the competing decisions are indistinct, or the method cannot observe the mechanism.

Do not borrow trust from the participant

Fake doors and Wizard of Oz setups can involve deception. Convenience does not make that acceptable.

Explain the research purpose, what will happen, data collection and use, observers, recording, retention, withdrawal, and the prototype’s limits in language participants can understand.

GOV.UK’s informed-consent guidance covers the consent and data-handling elements of that list for government research. It does not address prototype fidelity.

If revealing a simulation beforehand would invalidate the task, obtain appropriate ethical and legal review. Minimise the deception, prevent consequential actions, and debrief participants promptly.

Some domains should not use the method at all.

Never accept a real payment, expose a real account, imply clinical or financial certainty, or trigger an operational commitment merely to make a prototype feel authentic.

Protect shared prototypes from accidental discovery. Use synthetic or minimised data where possible, restrict access, remove credentials, define retention, and make the non-production status unmistakable.

Close the prototype deliberately

Prototypes accumulate debt when their shortcuts stop being visible. Hard-coded data becomes a “temporary” database. A manual approval is mistaken for automation. Unsecured code is copied because it already looks complete.

End the work with a transition contract:

  • the decision made and the evidence boundary;
  • what remains unknown and who owns it;
  • which behaviour, content, and requirements should be retained;
  • whether code, data, accounts, and environments will be destroyed or quarantined;
  • which production checks must start from first principles;
  • an owner and date for disposal or review.

An MVP is not a promoted prototype. It is a live product exposure with real obligations for reliability, privacy, support, measurement, and recovery.

Nor is a polished prototype a design handoff. A handoff must transfer decisions, states, constraints, ownership, and unresolved risks.

Screens alone preserve the most misleading part of the artifact.

A fictional exception-routing test

Consider an explicitly fictional procurement product. The team must choose between automatic routing of supplier exceptions and a guided review that keeps an operator responsible.

The uncertainty is whether operators can detect a wrong category, understand its consequence, and correct it before routing. It is not yet whether a classifier can reach production accuracy.

The team creates a Wizard of Oz prototype. A researcher supplies suggested categories behind the interface. The task uses synthetic cases shaped like the real work, including ambiguity and missing documents.

The critical realism boundary includes terminology, permission levels, correction, and the downstream destination. Model latency, model quality, production integrations, and operational cost are explicitly simulated or excluded.

The team recruits operators who make comparable decisions and tests in their working environment. It records behaviour separately from interpretation and defines stop, revise, advance, and abandon conditions before sessions begin.

No result is claimed here. The example only shows how a prototype can examine the human decision while refusing to imply that the automation is feasible or desirable.

The prototype’s best outcome is a clear decision

Speed comes from refusing to build qualities that cannot affect the current choice. Credibility comes from refusing to omit the qualities that can.

A prototype may earn another test, a technical investigation, a production-quality build, or disposal. All four can be valid outcomes.

The waste begins when the artifact survives but its claim, evidence boundary, and decision disappear.

Sources

Related books

If you want to go further on this topic, these are two good places to start.

02

product

The Lean Startup

by Eric Ries

How today's entrepreneurs use continuous innovation to create radically successful businesses, introducing Build-Measure-Learn and validated learning.

Some outbound links are affiliate links and support independent bookstores.