A shared vocabulary for reporting AI failures without collapsing them into “bugs”
What this artifact is
A reporting and editorial typology that helps journalists classify AI-related incidents accurately, before analysis or framing decisions are made.
Its purpose is to prevent a common failure mode in coverage:
treating fundamentally different kinds of AI harm as variations of the same technical problem.
This is not an academic taxonomy. It is a working classification tool designed to sit:
- in a reporter’s notebook,
- in an investigations desk playbook,
- or as an internal reference alongside standards & practices.
Why this artifact is necessary
One reason AI coverage repeatedly misfires is that journalists lack a shared incident vocabulary. As a result:
- governance failures are reported as software defects,
- deployment decisions are described as “misuse,”
- authority gaps are narrated as surprises,
- and systemic risks are individualized.
This typology gives newsrooms a way to ask, early:
“What kind of failure is this, actually?”
Core Principle
Not all AI failures are alike.
They differ in:
- cause,
- locus of responsibility,
- appropriate remedies,
- and public significance.
Misclassification leads directly to misframing.
The Typology (Five Incident Types)
Each AI-related incident should be classified under one or more of the following categories. Multiple categories may apply.
Type 1: Product / Model Failure
Definition:
The system behaves in a way that contradicts its intended technical specifications.
Characteristics
- Incorrect outputs relative to known constraints
- Regression after an update
- Performance degradation
- Hallucinations in contexts where accuracy was expected
Typical signals
- “Tests show accuracy dropped”
- “Model gave wrong factual answer”
- “Update worsened performance”
What journalists should ask
- Was this behavior unexpected by designers?
- Would fixing the model remove the harm?
What this type does not explain
- Why the system was trusted
- Why outputs were acted upon
- Why harm occurred at scale
Type 2: Deployment Failure
Definition:
The system functions as designed, but is deployed in an inappropriate context.
Characteristics
- High-risk domain (health, law, welfare)
- Inadequate safeguards for context
- Mismatch between system capability and use case
Typical signals
- “Tool not designed for this purpose”
- “Warnings existed but were ignored”
- “Used beyond original scope”
What journalists should ask
- Who approved this deployment?
- What risk assessment preceded it?
Type 3: Interface-Governed Failure
Definition:
Harm arises from how outputs are presented, not what they contain.
Characteristics
- Single-answer presentation
- Confident or authoritative tone
- Default visibility
- Summaries replacing primary sources
Typical signals
- “Users assumed the AI was authoritative”
- “People relied on summaries”
- “Output positioned as sufficient”
What journalists should ask
- How did the interface shape interpretation?
- Was sufficiency implied?
Type 4: Governance Failure
Definition:
No clear authority authorized, constrained, or monitored the system’s use.
Characteristics
- No accountable decision-maker
- Responsibility diffused across actors
- Safeguards promised but unenforced
- Oversight reactive rather than binding
Typical signals
- “No one appears responsible”
- “Review announced, system continues”
- “Policies exist, enforcement unclear”
What journalists should ask
- Where does authority reside?
- What mechanisms bind behavior?
Type 5: Responsibility Displacement Failure
Definition:
Responsibility for harm is assigned to actors without control over the system.
Characteristics
- Users blamed for reliance
- Professionals held accountable for AI-mediated decisions
- Moderators absorbing harm without power
Typical signals
- “Users should have known better”
- “Humans remain responsible”
- “Frontline workers blamed”
What journalists should ask
- Do these actors have authority to change system behavior?
- If not, why are they responsible?
How to Use the Typology (Workflow)
Before drafting:
- Identify the incident.
- Assign at least one incident type.
- Check whether multiple types apply.
- Use the dominant type to guide framing.
Editorial red flag:
If a story is framed as Type 1 (product failure) but is clearly Type 3–5, revision is required.
Example (from the Guardian corpus)
Incident: Google AI Overviews giving misleading health information.
- Product failure? ☐ (outputs often factually plausible)
- Deployment failure? ☑ (used in health context)
- Interface-governed failure? ☑ (single authoritative summaries)
- Governance failure? ☑ (no clear authorization regime)
- Responsibility displacement? ☑ (users told to verify)
Conclusion:
This is not primarily a bug story.
What This Typology Prevents
- “Just a glitch” narratives
- Over-reliance on accuracy metrics
- Individual blame for structural failures
- Endless calls for “better models” where governance is missing
What This Typology Enables
- Clearer investigative angles
- More precise questioning of institutions
- Better headline and lede framing
- Stronger accountability journalism
Intended Users
- Reporters (technology, health, investigations)
- Editors
- Journalism educators
- Media watchdogs
Member discussion: