A shared vocabulary for reporting AI failures without collapsing them into “bugs”

What this artifact is

A reporting and editorial typology that helps journalists classify AI-related incidents accurately, before analysis or framing decisions are made.

Its purpose is to prevent a common failure mode in coverage:

treating fundamentally different kinds of AI harm as variations of the same technical problem.

This is not an academic taxonomy. It is a working classification tool designed to sit:

  • in a reporter’s notebook,
  • in an investigations desk playbook,
  • or as an internal reference alongside standards & practices.

Why this artifact is necessary

One reason AI coverage repeatedly misfires is that journalists lack a shared incident vocabulary. As a result:

  • governance failures are reported as software defects,
  • deployment decisions are described as “misuse,”
  • authority gaps are narrated as surprises,
  • and systemic risks are individualized.

This typology gives newsrooms a way to ask, early:

“What kind of failure is this, actually?”

Core Principle

Not all AI failures are alike.
They differ in:

  • cause,
  • locus of responsibility,
  • appropriate remedies,
  • and public significance.

Misclassification leads directly to misframing.


The Typology (Five Incident Types)

Each AI-related incident should be classified under one or more of the following categories. Multiple categories may apply.


Type 1: Product / Model Failure

Definition:
The system behaves in a way that contradicts its intended technical specifications.

Characteristics

  • Incorrect outputs relative to known constraints
  • Regression after an update
  • Performance degradation
  • Hallucinations in contexts where accuracy was expected

Typical signals

  • “Tests show accuracy dropped”
  • “Model gave wrong factual answer”
  • “Update worsened performance”

What journalists should ask

  • Was this behavior unexpected by designers?
  • Would fixing the model remove the harm?

What this type does not explain

  • Why the system was trusted
  • Why outputs were acted upon
  • Why harm occurred at scale

Type 2: Deployment Failure

Definition:
The system functions as designed, but is deployed in an inappropriate context.

Characteristics

  • High-risk domain (health, law, welfare)
  • Inadequate safeguards for context
  • Mismatch between system capability and use case

Typical signals

  • “Tool not designed for this purpose”
  • “Warnings existed but were ignored”
  • “Used beyond original scope”

What journalists should ask

  • Who approved this deployment?
  • What risk assessment preceded it?

Type 3: Interface-Governed Failure

Definition:
Harm arises from how outputs are presented, not what they contain.

Characteristics

  • Single-answer presentation
  • Confident or authoritative tone
  • Default visibility
  • Summaries replacing primary sources

Typical signals

  • “Users assumed the AI was authoritative”
  • “People relied on summaries”
  • “Output positioned as sufficient”

What journalists should ask

  • How did the interface shape interpretation?
  • Was sufficiency implied?

Type 4: Governance Failure

Definition:
No clear authority authorized, constrained, or monitored the system’s use.

Characteristics

  • No accountable decision-maker
  • Responsibility diffused across actors
  • Safeguards promised but unenforced
  • Oversight reactive rather than binding

Typical signals

  • “No one appears responsible”
  • “Review announced, system continues”
  • “Policies exist, enforcement unclear”

What journalists should ask

  • Where does authority reside?
  • What mechanisms bind behavior?

Type 5: Responsibility Displacement Failure

Definition:
Responsibility for harm is assigned to actors without control over the system.

Characteristics

  • Users blamed for reliance
  • Professionals held accountable for AI-mediated decisions
  • Moderators absorbing harm without power

Typical signals

  • “Users should have known better”
  • “Humans remain responsible”
  • “Frontline workers blamed”

What journalists should ask

  • Do these actors have authority to change system behavior?
  • If not, why are they responsible?

How to Use the Typology (Workflow)

Before drafting:

  1. Identify the incident.
  2. Assign at least one incident type.
  3. Check whether multiple types apply.
  4. Use the dominant type to guide framing.

Editorial red flag:
If a story is framed as Type 1 (product failure) but is clearly Type 3–5, revision is required.


Example (from the Guardian corpus)

Incident: Google AI Overviews giving misleading health information.

  • Product failure? ☐ (outputs often factually plausible)
  • Deployment failure? ☑ (used in health context)
  • Interface-governed failure? ☑ (single authoritative summaries)
  • Governance failure? ☑ (no clear authorization regime)
  • Responsibility displacement? ☑ (users told to verify)

Conclusion:
This is not primarily a bug story.


What This Typology Prevents

  • “Just a glitch” narratives
  • Over-reliance on accuracy metrics
  • Individual blame for structural failures
  • Endless calls for “better models” where governance is missing

What This Typology Enables

  • Clearer investigative angles
  • More precise questioning of institutions
  • Better headline and lede framing
  • Stronger accountability journalism

Intended Users

  • Reporters (technology, health, investigations)
  • Editors
  • Journalism educators
  • Media watchdogs