Practical AI Governance: Build a Decision and Evidence System
Govern AI across its lifecycle with clear decision rights, scoped use, evidence, release gates, supplier controls, incident response, and retirement.
Piotr Ciechowicz
Product manager · developer
Updated July 13, 2026
On this page15 sections
- 01Govern a use case, not a technology label
- 02Assign rights to decide, not merely roles to attend
- 03Classify consequence before selecting controls
- 04Define intended purpose and excluded use
- 05Build an assurance case around the release claim
- 06Preserve provenance across the whole system
- 07Make human control a real operating capability
- 08Use release and change gates
- 09Treat suppliers as moving product dependencies
- 10Monitor consequence, change and control failure
- 11A fictional change review
- 12Retire the system, not only the interface
- 13The AI governance decision record
- 14Sources
- 15Read next
Suppose a model has passed its evaluation. The privacy review is complete. The launch decision is recorded.
Two weeks later, the supplier replaces the model behind the same API name. The assistant begins refusing ordinary requests, while a few unsupported answers become harder for reviewers to spot.
Which evidence is now stale? Who can pause the feature? Does the change require a new release decision, or can engineering treat it as routine maintenance?
An AI policy will not answer those questions. A register of approved tools will not answer them either.
AI governance is the operating system that connects changing model behaviour to product consequences. It assigns decision rights, states what evidence is sufficient, and defines when a decision must be reopened.
The thing to govern is not “the model”. It is a use case across its full lifecycle: purpose, data, interface, operator, affected people, supplier, monitoring, incidents, and retirement.
Govern a use case, not a technology label
“We use a language model” says little about consequence.
The same model might help an employee rewrite a heading, rank customer complaints, or decide which account receives a fraud hold. Its role, data and downstream authority make those uses materially different.
Create one inventory record for each use case, including experiments that touch real data or influence real work. Record at least:
- the user, affected people, and decision or action;
- the intended purpose and business owner;
- the model’s role: draft, suggest, rank, decide, or act;
- inputs, outputs, data sensitivity, and permitted data use;
- integrations and downstream effects;
- expected benefit and plausible failure consequences;
- human control, correction, contest and fallback routes;
- model, vendor and internal system owners;
- current lifecycle state and next review trigger.
One system may need several records. Drafting replies and routing urgent cases have different consequences, so they are separate use cases.
The AI product strategy guide covers the upstream choice of where AI could create value. Governance begins when that bet needs accountable boundaries and evidence.
Assign rights to decide, not merely roles to attend
Cross-functional participation can still leave the decisive question ownerless.
For each use case, name who can make five different decisions:
- Propose: define the purpose, audience and expected value.
- Challenge: test assumptions from technical, security, data, legal, domain and affected-person perspectives.
- Accept: approve the bounded use and its residual exposure.
- Operate: monitor behaviour, investigate signals and maintain controls.
- Pause or retire: stop the use when evidence, context or obligations change.
The same person may hold several rights in a small team. The rights must still be explicit.
Specialists retain authority that comes from law, policy or professional responsibility. A product owner cannot approve non-compliance, and a committee vote cannot turn missing evidence into adequate evidence.
Give one person ownership of the release decision. Record dissent and unresolved uncertainty for that person to see.
Classify consequence before selecting controls
A three-colour “AI risk” label is too coarse to design a release process.
Describe the consequence through several dimensions:
- who can be affected, including people who never use the product;
- what the system can change, recommend, reveal, delay or deny;
- severity, scale, duration and reversibility of a wrong outcome;
- how likely an error is to be noticed before it propagates;
- the system’s autonomy and the real authority of a reviewer;
- data sensitivity, provenance and permitted purpose;
- variation across tasks, languages, groups and operating contexts;
- dependence on a supplier, model or data source the team cannot inspect or control.
This description should change the controls. A reversible drafting aid may need sampled review and a reliable fallback. A system influencing access, employment, health, safety or money may need specialist review and much stronger evidence.
The NIST AI Risk Management Framework organises risk work as Govern, Map, Measure and Manage across the lifecycle.
NIST describes AI RMF 1.0 as voluntary, rights-preserving, non-sector-specific and use-case agnostic. Its actions are not an ordered checklist, and NIST is currently revising the framework.
For generative AI, the NIST Generative AI Profile is a cross-sector companion to AI RMF 1.0.
These resources help a team structure questions and actions. Applying them does not establish that a system is safe, fair or compliant.
Define intended purpose and excluded use
“Improve productivity” cannot govern a product. It does not constrain input, output, audience or action.
Write the intended purpose as an observable product behaviour:
Draft an internal incident summary from the selected support records so a trained agent can verify the evidence before escalating the case.
Then state excluded uses with equal precision:
- may not send a message or update an account;
- may not infer customer intent beyond the cited records;
- may not process records outside the approved support workspace;
- may not be used to assess staff performance;
- may not replace the escalation owner’s decision.
Excluded use belongs in access controls, interface language, evaluation cases, training and monitoring. A sentence in a policy cannot stop a convenient repurposing.
Reopen the use-case decision when a team adds a new audience, data source, action, integration or purpose. “The model has not changed” is irrelevant if the consequence has.
Build an assurance case around the release claim
A passing test suite is evidence about tests. It is not proof that the use case is ready.
For a consequential release, write a short assurance case:
- Claim: what the team believes the bounded system can do acceptably.
- Evidence: observations that support the claim in the intended context.
- Coverage: users, tasks, conditions and failure modes represented.
- Limit: what the evidence cannot establish.
- Control: what prevents or reduces an unacceptable consequence.
- Residual exposure: what can still go wrong and who may bear it.
- Owner: who accepts the claim for this scope.
- Reopen trigger: which change or signal invalidates the decision.
This is a practical product pattern, not a claim that one document satisfies a standard or law.
Evidence may include data-lineage review, task-specific evaluation, difficult-case testing, human-factors testing, security review, recovery exercises, accessibility checks and observation of the complete workflow.
A single average score rarely supports the whole claim. Evidence should reveal where performance differs and whether the safeguards work when the model does not.
Preserve provenance across the whole system
An output may depend on more than a model version.
Record the model or endpoint, provider, system instructions, prompt template, retrieval index, tools, policies, thresholds, input sources and relevant interface version.
For data, record origin, owner, permitted purpose, transformations, quality limits, retention and deletion behaviour. For external models, record the supplier’s stated data handling and the contract version reviewed.
Link each release to the evaluation set, results, known limitations, approval and rollback option. The goal is to reconstruct the product behaviour that was assessed, not to collect documents nobody can use.
Provenance also exposes blind spots. A supplier may not reveal training data or may reserve the right to change a hosted model. That uncertainty belongs in the assurance case and supplier decision.
Make human control a real operating capability
“Human in the loop” often names a location in a diagram, not an effective control.
The person must have enough time, competence, evidence, authority and interface support to notice a problem and choose another action. If rejecting every output threatens their target, formal authority may not survive operational pressure.
Test human control with difficult cases. Can the reviewer find the decisive evidence? Can they override the result without a workaround? Does the system preserve the challenge and the final reason?
People affected by a consequential output may also need a route to understand, correct or contest it. The AI UX guide covers that interaction in depth.
The ICO’s AI and data protection guidance is narrower than general AI governance.
It explains the ICO’s interpretation and recommended practice for AI processing personal data under the UK data-protection regime. It is not a statutory code or an exhaustive guide, and it does not govern uses outside that remit.
Where UK personal data is in scope, involve the appropriate privacy specialist to determine obligations. This article is not legal advice.
Use release and change gates
Governance becomes useful when evidence can change a commitment.
A release gate should allow four outcomes: approve the bounded use, approve with conditions, request more evidence, or reject it.
Name the exact model, configuration, data, users, authority and fallback being approved. Define what counts as a material change.
Common reopen triggers include:
- a new model, model version, provider or hosting arrangement;
- changes to prompts, tools, retrieval, policy or thresholds;
- new input data, audience, language, market or downstream action;
- a supplier’s terms, retention practice or safety control changing;
- performance moving outside the evaluated range;
- a severe incident, repeated contest, or newly identified affected group;
- a change in applicable law, official guidance or internal risk appetite.
Not every edit needs a board. Route a change by its possible consequence and the evidence it invalidates. Demand more when exposure expands.
Treat suppliers as moving product dependencies
A vendor assessment made during procurement becomes stale if the service changes.
Ask what the supplier will disclose about model updates, incidents, data location, retention, subprocessors, evaluation, security controls and service changes.
Determine whether versions can be pinned and how long an old version remains available.
Contractual questions need commercial and legal owners. Product still needs an operating answer: how will the team detect a change, retest affected claims, communicate a changed limitation and restore a fallback?
The UK AI Playbook covers lifecycle management, traceability and supply-chain responsibility.
It is public-sector guidance for UK government organisations, not a universal standard. Its lifecycle and supplier questions are still useful prompts for a product team to adapt to its own context.
Do not outsource the release decision to a vendor’s trust page. Their evidence may support one part of your case; it cannot assess your use, interface, operators or affected people.
Monitor consequence, change and control failure
Model quality is only one production signal.
Monitor the inputs and contexts reaching the system, unsupported or abstained outputs, user corrections, overrides, contests, escalations, downstream actions and recovery attempts.
Segment signals where aggregate performance could hide a meaningful failure. Track whether operators have enough time to review, whether fallback paths work, and whether excluded uses appear in practice.
Drift can enter through data, user behaviour, policy, workflow or supplier changes. A stable offline score does not show that the product is still being used for the approved purpose.
Define an incident before one happens. Include criteria for severity, containment authority, evidence preservation, specialist escalation, supplier contact, communication to affected people, and the decision to resume, narrow or retire.
An incident review should update the use-case record and assurance case. Closing a ticket without reopening the governing claim leaves the system ready to repeat the same mistake.
A fictional change review
Consider a fictional customer-support product that drafts internal incident summaries. It cites selected records, cannot contact a customer, and requires an agent to accept each claim before escalation.
The supplier announces a model upgrade. The product team does not ask only whether the new model is “better”. It maps the change to the approved claims.
Replay tests show that citation format remains valid, but the model now combines contradictory account records into one fluent statement. Reviewers can see both sources, yet the interface does not mark the conflict.
The release owner rejects the upgrade for the current configuration. Engineering pins the previous version while the team adds conflict detection and an abstention case to the evaluation set.
The supplier cannot guarantee long-term access to the old version. That new dependency risk is recorded with a migration deadline and a manual-summary fallback.
No measured improvement or real company outcome is claimed here. The example is fictional. It shows how provenance, evidence and authority turn a supplier notice into a product decision.
Retire the system, not only the interface
Every governed use case needs an exit route.
Retirement may require disabling integrations and credentials, removing user access, preserving required decision records, applying retention and deletion rules, notifying operators, and restoring a non-AI path.
Check downstream dependencies before shutdown. A team may have copied the model’s output into reports, rules or another product, even after the original feature disappears.
The governance record should close with the reason, owner, effective date, remaining data, unresolved incidents and any continuing obligations.
The AI governance decision record
Keep one durable record that links, rather than duplicates, the evidence:
- use case, intended purpose and excluded use;
- affected people, action, consequence and reversibility;
- applicable jurisdictions, policies and specialist decisions;
- data, model, supplier, configuration and provenance;
- decision rights, escalation and pause authority;
- assurance claim, evidence, coverage and limits;
- human control, contest, fallback and incident routes;
- release outcome, residual exposure and rationale;
- monitoring signals, thresholds and named responders;
- material-change, review and retirement triggers.
The record is not a compliance certificate. It is a way to make the decision inspectable while it is still possible to challenge, narrow or stop it.
Regulation adds context-specific duties to this operating system. It does not replace it.
For teams within the EU AI Act’s scope, start with the official Regulation text.
Check the Commission’s current application timeline separately because implementation dates are changing.
As checked on 13 July 2026, the Commission says the Act applies in phases and becomes broadly applicable on 2 August 2026, with exceptions.
Its current page also reports a political agreement to move specified high-risk rules to December 2027 and rules for AI embedded in regulated products to August 2028.
Applicability depends on jurisdiction, system, role and use. Verify the current legal text and enacted amendments with qualified counsel rather than turning this summary into a compliance conclusion.
Good governance keeps the decision alive. When behaviour, evidence, context or consequence changes, the team knows what must be proved again and who can act.
Sources
- AI Risk Management Framework 1.0 - NIST
- Generative Artificial Intelligence Profile - NIST
- Regulation (EU) 2024/1689 - EUR-Lex
- AI Act application timeline - European Commission
- Artificial Intelligence Playbook for the UK Government - GOV.UK
- Guidance on AI and data protection - UK Information Commissioner’s Office
Read next
AI-Assisted Workflows: Design Work That Can Be Checked applies these controls to the work product, context, review burden and hand-offs between a model and people.
NLP Product Workflows: From Model Output to Operable Capability shows how the same controls become acceptance, fallback, monitoring, and ownership decisions in an NLP workflow.
Related books
Two books to
read next.
If you want to go further on this topic, these are two good places to start.
01
data
Lean Analytics
by Alistair Croll & Benjamin Yoskovitz
How to use data to build a better startup faster, with frameworks for identifying the right metrics at each stage of company growth.
02
technology
Natural Language Processing with Transformers
by Lewis Tunstall, Leandro von Werra & Thomas Wolf
Building language applications with Hugging Face, covering modern NLP architecture from the creators of the Transformers library.
Some outbound links are affiliate links and support independent bookstores.