This is checklist is for ministries, regulators, hospitals, councils, universities, platforms, or large organizations assessing their own AI use or an incident involving a vendor system.
Purpose:
To determine whether an AI-related incident reflects:
- a localized operational error, or
- a systemic governance failure requiring structural intervention.
Each item is Yes / No / Unknown.
“Unknown” is a finding: it signals loss of institutional visibility.
I. Authority & Mandate
- Is there a documented mandate specifying what epistemic role this system is allowed to play (e.g. advisory vs authoritative)?
☐ Yes ☐ No ☐ Unknown - Are there explicit boundaries defining where the system must not provide conclusive outputs (e.g. health, legal, eligibility decisions)?
☐ Yes ☐ No ☐ Unknown - Is a specific institutional role or unit accountable for epistemic harm caused by this system’s outputs?
☐ Yes ☐ No ☐ Unknown - Do staff understand whether they are expected to defer to, challenge, or merely consider the system’s outputs?
☐ Yes ☐ No ☐ Unknown
II. Interface & Use Context
- Does the system present single synthesized outputs by default in high-stakes contexts?
☐ Yes ☐ No ☐ Unknown - Are uncertainty, disagreement, or evidence strength visible to users at the point of decision?
☐ Yes ☐ No ☐ Unknown - Does the interface nudge users toward acceptance (speed, tone, placement) rather than reflection?
☐ Yes ☐ No ☐ Unknown - Have interface design choices been reviewed as governance decisions, not just UX choices?
☐ Yes ☐ No ☐ Unknown
III. Compression & Information Loss
- Is the system used to summarize, triage, or condense complex case material?
☐ Yes ☐ No ☐ Unknown - Are there safeguards to prevent loss of urgency, severity, or minority cases during summarization?
☐ Yes ☐ No ☐ Unknown - Is there a documented process for users to inspect what was omitted or downweighted?
☐ Yes ☐ No ☐ Unknown - Have false reassurance or premature closure risks been explicitly assessed?
☐ Yes ☐ No ☐ Unknown
IV. Bias & Differential Impact
- Has the system been tested for differential effects across populations relevant to its use (e.g. gender, age, vulnerability)?
☐ Yes ☐ No ☐ Unknown - Are bias assessments model-specific and task-specific (not generic assurances)?
☐ Yes ☐ No ☐ Unknown - Is bias monitoring ongoing rather than one-time?
☐ Yes ☐ No ☐ Unknown - Is there a mechanism to surface harms that arise through omission or framing, not just explicit error?
☐ Yes ☐ No ☐ Unknown
V. Institutional Comprehension & Control
- Can the institution clearly identify the model(s) in use, including versioning and update cadence?
☐ Yes ☐ No ☐ Unknown - Is there internal capacity to audit or independently evaluate system behavior?
☐ Yes ☐ No ☐ Unknown - Are changes in model behavior detectable by users or oversight staff?
☐ Yes ☐ No ☐ Unknown - Was reliance on the system introduced before governance and oversight structures were in place?
☐ Yes ☐ No ☐ Unknown
VI. Governance & Enforcement
- Are safeguards enforceable at the same speed and scale as deployment?
☐ Yes ☐ No ☐ Unknown - Do public statements or policies correspond to operational constraints?
☐ Yes ☐ No ☐ Unknown - Is there a defined escalation pathway when harm is identified (pause, rollback, containment)?
☐ Yes ☐ No ☐ Unknown - Does delay in review or regulation effectively allow continued operation?
☐ Yes ☐ No ☐ Unknown
VII. Scale, Speed, and Irreversibility
- Does the system operate or propagate at a scale that exceeds the institution’s corrective capacity?
☐ Yes ☐ No ☐ Unknown - Are retraction, correction, or containment mechanisms slower than dissemination?
☐ Yes ☐ No ☐ Unknown - Have downstream effects persisted after partial fixes or public acknowledgment?
☐ Yes ☐ No ☐ Unknown - Is full reversal of harm realistically achievable?
☐ Yes ☐ No ☐ Unknown
Interpretation Guide (Internal Use)
- Multiple “Unknown” responses → loss of institutional visibility
- Clusters of “Yes” in I–II → authority/interface risk
- Clusters in III–IV → harm and equity risk
- Clusters in V–VII → systemic governance failure
A critical signal is misalignment:
- High reliance + low comprehension
- Strong statements + weak enforcement
- Large scale + slow correction
What this checklist is designed to do
- Force explicit answers where ambiguity is often tolerated
- Distinguish operational risk from structural governance failure
- Create an internal record of what the institution knew, when
It is not a compliance box-ticker.
It is a visibility and accountability instrument.
Member discussion: