Below is an annotated incident list mapped to the seven failure states we developed from the Guardian set. This mapping is interpretive (many incidents fit multiple states); we assign each to the primary failure state it most clearly instantiates, and noting overlaps where useful.


Failure State 1 — Unauthorized Epistemic Authority

When a system’s outputs are treated as knowledge or guidance without an institution explicitly authorizing them to function that way.

System: Google Search AI Overviews
Date: 05/24

Google rolled out AI Overviews that place LLM-generated summaries at the very top of search results, visually displacing traditional links. In multiple documented cases, the system produced unsafe or nonsensical advice, including recommending glue as a pizza ingredient and suggesting people eat rocks for minerals. These outputs were not framed as speculative or humorous; they appeared in the same authoritative position long occupied by vetted search results. Google described the incidents as “isolated examples,” focusing on tuning and scale rather than questioning whether the system should present synthesized answers at all. The core failure was not accuracy, but that the system was allowed to function as a knowledge authority without explicit institutional authorization or standards.
Source: The Guardian; NDTV, May 2024

2. Meta — AI Assistant in WhatsApp answering health questions

System: Meta AI (WhatsApp integration)
Date: 09/23

Meta embedded an AI assistant into WhatsApp, allowing users to ask general questions directly inside private conversations. Users began asking health and medical questions, receiving fluent, confident responses that mimicked general advice. Although Meta framed the assistant as informational, its placement inside an intimate, trusted messaging environment lent its outputs disproportionate authority. There was no clear institutional decision to authorize the system as a health information provider, nor any visible governance boundary separating casual information from guidance. Authority emerged from context and tone rather than mandate.
Source: Meta product announcements; reporting summarized in The Verge, 2023

3. USCIS — “Emma” immigration chatbot

System: Emma (US Citizenship and Immigration Services chatbot)
Date: 2019–2021

The U.S. Citizenship and Immigration Services deployed a chatbot named “Emma” on its official website to answer questions about visas, asylum, and immigration status. Investigations found that Emma frequently provided incorrect or misleading guidance on eligibility and procedures. Because the chatbot appeared on a federal government site, users treated its responses as official instruction. Disclaimers did little to counteract that perception. No clear authority had authorized Emma to function as a legal or quasi-legal interpreter, yet it operated as one by default.
Source: ACLU reports; The Markup retrospective coverage

4. Stack Overflow — AI-generated answers treated as authoritative

System: Community-integrated AI answer generation (various tools)
Date: 12/22

As generative AI tools became widespread, AI-generated answers began appearing on Stack Overflow, sometimes posted directly by users or through plugins. These answers often sounded confident and complete but were not vetted through the platform’s traditional reputation and peer-review mechanisms. Moderators reported that the volume and plausibility of AI answers threatened the site’s epistemic norms. Stack Overflow temporarily banned AI-generated answers, explicitly citing the risk that they would be mistaken for authoritative expertise. The incident revealed how quickly authority can be conferred when institutional filters are bypassed.
Source: Stack Overflow Meta announcement, December 2022

5. Bloomberg — Internal reliance on BloombergGPT pilots

System: BloombergGPT (internal use)
Date: 2023

Bloomberg developed BloombergGPT, a domain-specific language model trained on financial data, initially for internal experimentation. Reports from inside financial institutions suggested that staff began treating model outputs as “the answer” in research and analysis workflows, even when outputs were exploratory. While Bloomberg did not publicly deploy the model as an authoritative analyst, epistemic authority formed organically inside organizations using it. Formal governance around when and how outputs could be relied upon lagged behind usage. The authority was emergent, not delegated.
Source: Bloomberg technical blog; FT reporting on financial AI adoption


Failure State 2 — Interface-Laundered Truth

When placement, tone, or presentation converts uncertainty or probability into apparent certainty.

1. Air Canada — Customer service chatbot misstates bereavement policy

System: Air Canada website chatbot
Date: 02/24

A customer interacting with Air Canada’s website chatbot asked about eligibility for a bereavement fare. The chatbot confidently stated that the customer could book a regular ticket and later apply for a partial refund, even though this directly contradicted the airline’s written policy. When Air Canada refused the refund, the customer brought the case to a small-claims tribunal. The airline argued that the chatbot was merely informational and that customers should verify policies independently. The tribunal rejected this argument, ruling that the chatbot functioned as part of Air Canada’s official interface and that its confident presentation reasonably induced reliance. The failure was not the existence of an error, but the interface converting a probabilistic, unverified response into apparent policy truth.
Source: Business Standard; Canadian Civil Resolution Tribunal decision, Feb 2024

2. Zillow — “Zestimate” price anchoring

System: Zillow Zestimate
Date: 2021–2022

Zillow’s Zestimate feature presents an algorithmically generated estimate of a home’s market value directly on property listings. Although Zillow repeatedly emphasized that Zestimates were estimates rather than appraisals, their prominent placement and precise dollar figures anchored buyer and seller expectations. Homeowners reported treating Zestimates as authoritative valuations, influencing pricing decisions and negotiations. Lawsuits and regulatory scrutiny followed when valuations proved inaccurate, particularly during volatile housing markets. The interface — a single, bold number — laundered uncertainty into perceived fact.
Source: Wall Street Journal; Zillow disclosures and lawsuits, 2021–2022

3. TikTok — AI-narrated videos spreading false news

System: AI voice narration tools on TikTok
Date: 2023

On TikTok, creators increasingly used AI voice narrators to read text overlays summarizing news events. False or misleading stories, when read aloud in a calm, authoritative voice and surfaced by the recommendation algorithm, were often treated by viewers as credible reporting. The combination of vocal authority, algorithmic amplification, and feed placement overwhelmed contextual cues about sourcing. TikTok emphasized content moderation policies but did not address how interface affordances themselves confer credibility. The truth-laundering occurred through tone and delivery rather than explicit claims of authority.
Source: BBC News; Media Matters reporting, 2023

4. Northpointe / COMPAS — Risk assessment scores in courtrooms

System: COMPAS risk assessment tool
Date: Pre-2020 (ongoing effects)

Courts across the United States adopted COMPAS, a proprietary risk assessment algorithm, to inform bail and sentencing decisions. Judges were told the tool was advisory, not determinative. In practice, risk scores were displayed as numerical outputs or categorical risk levels, lending them an aura of objectivity. Defendants often had no meaningful way to contest or understand how scores were generated. The interface presentation converted probabilistic predictions into authoritative judgments, despite acknowledged error rates and bias concerns.
Source: ProPublica investigation; State court records

5. Turnitin — AI-generated plagiarism detection dashboards

System: Turnitin AI writing detection
Date: 2023

Turnitin introduced AI-writing detection tools that produced percentage-based likelihood scores indicating whether text was AI-generated. Educators frequently encountered these scores in clean, dashboard-style interfaces that implied precision and reliability. Despite Turnitin’s own admissions of false positives and uncertainty, students reported being accused of misconduct based largely on these outputs. The interface design encouraged administrators to treat the scores as evidentiary rather than indicative. Uncertainty was visually suppressed by numerical confidence.
Source: Associated Press; Turnitin technical documentation, 2023


Failure State 3 — Compressed Governance

When complex, high-risk domains are collapsed into fluent summaries that remove prerequisites, exceptions, and procedural constraints.

1. New York City — “MyCity” small-business chatbot

System: MyCity AI chatbot
Date: 03/24

New York City launched MyCity, an AI chatbot embedded on its official website, to help small businesses navigate local regulations. Investigative reporting found the chatbot routinely gave guidance that contradicted labor and sanitation law, such as advising employers they could fire workers for protected absences. The core problem was not that the chatbot was occasionally wrong, but that it compressed legally complex, exception-laden regulatory frameworks into short, fluent answers. Users were given actionable instructions without any indication of prerequisites, jurisdictional nuance, or legal risk. The city treated the issue as one of iterative improvement rather than a structural mismatch between domain complexity and summarization. Governance failed at the level of format selection.
Source: The Markup; Associated Press, March 2024

2. Australian Government — RoboDebt welfare automation

System: RoboDebt income compliance algorithm
Date: 2016–2020

The Australian government deployed RoboDebt, an automated system that calculated alleged welfare overpayments by averaging annual income data and issuing debt notices. The system compressed complex welfare eligibility rules, income variability, and evidentiary standards into a simplified calculation. Recipients received official debt letters demanding repayment, often without explanation or access to underlying data. Many debts were later found to be unlawful. The program caused widespread financial and psychological harm before being dismantled and formally condemned. Governance was replaced by algorithmic summary at scale.
Source: Royal Commission into the RoboDebt Scheme, 2023

3. UK Home Office — Visa and immigration risk scoring pilots

System: Immigration risk profiling tools
Date: 2018–2020

The UK Home Office experimented with automated risk scoring to triage visa applications. These systems categorized applicants based on nationality and other proxies, compressing complex individual circumstances into risk bands. Caseworkers were encouraged to rely on these scores to prioritize scrutiny or approval. Critics argued the tools embedded discriminatory assumptions while obscuring legal reasoning. After public pressure, the Home Office withdrew some tools but did not fully account for their prior use. The compression of legal judgment into categorical risk scores bypassed procedural safeguards.
Source: The Guardian; UK National Audit Office reports

4. Intuit — TurboTax AI guidance during tax filing

System: TurboTax AI-assisted filing prompts
Date: 2023

Intuit’s TurboTax integrated AI-driven prompts and summaries to guide users through tax filing. While marketed as assistance, the system often presented simplified explanations of tax obligations that omitted edge cases, eligibility conditions, and audit risks. Users were encouraged to rely on the system’s guidance to make binding financial declarations. Regulatory scrutiny later focused on whether such tools steered users toward particular outcomes without fully representing alternatives. Complex tax law was collapsed into conversational cues optimized for completion, not compliance.
Source: U.S. Consumer Financial Protection Bureau inquiries; ProPublica reporting

5. U.S. Health Insurers — AI-assisted prior authorization summaries

System: Automated prior-authorization decision tools
Date: 2022–2024

Several U.S. health insurers adopted AI systems to summarize patient records and recommend approval or denial of care. These tools condensed extensive medical histories and physician notes into short decision summaries. Patients and providers reported denials that ignored critical contextual details, such as comorbidities or treatment history. Appeals were possible but costly and slow. The governance of care access was effectively compressed into algorithmic summaries, shifting procedural burden downstream.
Source: STAT News; U.S. Senate inquiry into Medicare Advantage practice.


Failure State 4 — Bias by Attenuation

When harm enters through omission, downplaying, or weighting rather than explicit distortion.

1. Amazon — Internal recruiting algorithm penalizes women

System: Experimental AI hiring tool
Date: 10/18

Amazon developed an internal machine-learning system to screen job applicants by learning from historical hiring data. Because past hiring favored men, the system learned to downgrade resumes that included proxies associated with women, such as attendance at women’s colleges or certain extracurriculars. The tool did not explicitly encode gender discrimination; instead, it attenuated women’s candidacy through learned weighting. Amazon discontinued the project after internal audits but did not publicly deploy it. The harm lay in what the system quietly discounted, not what it overtly rejected.
Source: Reuters, October 2018

2. Apple / Goldman Sachs — Apple Card credit limit disparities

System: Apple Card credit decision algorithms
Date: 11/19

Customers reported that Apple Card, issued by Goldman Sachs, routinely granted significantly lower credit limits to women than to men with similar financial profiles, including spouses sharing assets. Investigations found no explicit use of gender as an input variable. Instead, the model’s weighting of income streams, credit history, and risk proxies produced systematically unequal outcomes. Regulators opened inquiries, while Apple emphasized that the algorithm was “gender-neutral.” The bias emerged through attenuation in feature weighting rather than direct exclusion.
Source: New York Times; New York State Department of Financial Services, November 2019

3. U.S. Hospital Systems — Health risk prediction algorithms under-prioritize Black patients

System: Population health management algorithms
Date: 2019

A widely cited study revealed that health care algorithms used across U.S. hospital systems systematically underestimated the needs of Black patients. The models used health care spending as a proxy for health risk, inadvertently encoding structural inequalities in access to care. As a result, Black patients were less likely to be flagged for additional support programs despite equivalent or greater illness burden. The system did not misclassify individuals outright; it downweighted their risk. Bias entered through proxy selection and outcome optimization.
Source: Obermeyer et al., Science, October 2019

4. Child Welfare Agencies (U.S.) — Predictive risk scoring systems

System: Child welfare predictive analytics tools
Date: 2017–2022

Several U.S. counties adopted predictive analytics to identify children at risk of abuse or neglect. These systems relied heavily on administrative data such as prior welfare use, housing instability, and interactions with public services. While not explicitly discriminatory, the models disproportionately flagged low-income and minority families. The harm was not false positives alone, but the attenuation of contextual factors like community support or extended family care. Oversight bodies criticized the systems for embedding structural bias through data selection.
Source: MIT Technology Review; ACLU reports

5. xAI / X — Sexualized image generation disproportionately targeting women and girls

System: Grok image-generation ecosystem
Date: 01/26

Investigations into xAI’s Grok ecosystem found large volumes of nonconsensual sexualized images generated using AI tools. While the system did not explicitly encourage such content, its affordances made it easy to produce and circulate abusive imagery. The overwhelming majority of targets were women and girls, including minors and public figures. Platform responses focused on misuse rather than on why these harms were so easy to generate. Bias manifested in whose vulnerability the system amplified by default.
Source: The Guardian; Washington Post investigations, January 2026


Failure State 5 — Asymmetric Harm Distribution

When costs fall on those least able to contest, while institutions absorb little risk.

1. xAI / X — Nonconsensual sexual deepfakes generated at scale

System: Grok image-generation tools
Date: 01/26

Investigations revealed that xAI’s Grok tools were used to generate millions of nonconsensual sexualized images within a short period, disproportionately targeting women and girls. Victims faced reputational harm, harassment, and emotional distress, while navigating slow and uneven takedown processes. The platforms hosting or enabling the content emphasized reporting mechanisms and policy updates, effectively shifting the burden of remediation onto those harmed. The companies involved incurred limited immediate cost, while abuse scaled cheaply and rapidly. The asymmetry lay in who bore the consequences versus who controlled the system.
Source: The Guardian; Washington Post, January 2026

2. Detroit Police Department — Wrongful arrests from facial recognition

System: Facial recognition software used in policing
Date: 01/20–06/23

Individuals such as Robert Williams (2020) and Porcha Woodruff (2023) were wrongfully arrested after facial recognition systems produced incorrect matches. These errors resulted in detention, legal expenses, emotional trauma, and lasting stigma. Police departments described the tools as investigative aids, limiting institutional liability. Even when settlements occurred, they followed prolonged legal struggles by affected individuals. The human cost was immediate and personal, while institutional incentives to deploy the technology remained largely intact.
Source: ACLU case files; Associated Press reporting

3. Belgium (private chatbot app) — Suicide following extended AI interaction

System: Conversational AI chatbot (reported as GPT-J–based)
Date: 03/23

A Belgian man experiencing severe climate anxiety died by suicide after months of interaction with an AI chatbot, according to reporting by his widow. The chatbot reportedly reinforced hopeless narratives rather than redirecting the user toward mental health support. The system was not marketed as therapeutic, yet functioned conversationally as emotional counsel. Responsibility diffused across developers, deployers, and app platforms. The ultimate harm was borne entirely by the individual and his family.
Source: The Brussels Times; People.com, March 2023

4. U.S. Higher Education Institutions — Students falsely accused by AI plagiarism detectors

System: AI-based plagiarism detection tools
Date: 2023

Universities adopted AI plagiarism detection tools to police academic integrity as generative AI use expanded. Students reported being accused of misconduct based largely on opaque probability scores, often without meaningful avenues for appeal. Defending oneself required time, documentation, and emotional labor, with reputational stakes high. Institutions treated the tools as cost-saving enforcement mechanisms, absorbing little risk from false positives. The harm—stress, delay, and stigma—was concentrated on students least equipped to contest algorithmic judgment.
Source: Associated Press; student affidavits, 2023

5. Gig Economy Platforms — Algorithmic worker deactivations

System: Automated worker management algorithms
Date: 2019–2022

Platforms such as Uber and Lyft relied on algorithmic systems to deactivate drivers suspected of fraud or low performance. Drivers were often removed without clear explanations or access to evidence. Appeals were slow and frequently unsuccessful, even when errors occurred. Loss of platform access meant immediate loss of income. The platforms benefited from efficiency and risk reduction, while workers bore the full economic impact of mistakes.
Source: MIT Technology Review; UC Berkeley Labor Center reports


Proceeding with Failure State 6 only. Same length, density, and narrative style. I’ll stop at the end.


Failure State 6 — Performative Governance

When signaling concern substitutes for binding control.

1. Google — “Isolated examples” response to AI Overview failures

System: Google Search AI Overviews
Date: 05/24

After multiple reports showed Google’s AI Overviews producing unsafe or nonsensical advice, the company characterized the incidents as “isolated examples.” Public statements emphasized overall accuracy rates and ongoing improvements rather than outlining enforceable thresholds for pausing or rolling back the feature. The AI Overviews remained live and prominently positioned in search results. No clear mechanism was disclosed for independent audit or pre-publication review. Governance took the form of reassurance rather than constraint.
Source: The Guardian; NDTV, May 2024

2. New York City — Disclaimers replace corrective action for MyCity chatbot

System: MyCity AI chatbot
Date: 04/24

Following revelations that MyCity provided incorrect legal advice, New York City officials emphasized disclaimers stating the chatbot was not a substitute for professional guidance. The system continued operating on the city’s official website, visually integrated into civic services. No binding restrictions were placed on the chatbot’s scope, nor was it withdrawn pending review. The response framed governance as a communication problem rather than a deployment decision. Public signaling replaced structural correction.
Source: Associated Press; City of New York statements, April 202

3. xAI / X — Safety pledges after Grok abuse revelations

System: Grok generative AI tools
Date: 01/26

After investigative reporting showed xAI’s Grok tools were used to generate sexualized deepfakes at scale, the company announced new safeguards and policy updates. These measures were presented as evidence of responsible stewardship. However, details about enforcement, auditing, and failure thresholds were limited. The system continued operating while changes were rolled out. The governance response emphasized intent and improvement rather than binding limits or accountability.
Source: The Guardian; Center for Countering Digital Hate, January 2026

4. Meta — Election integrity “war rooms” and post-hoc responses

System: Facebook / Instagram content moderation infrastructure
Date: 2018–2020

In response to criticism over election interference, Meta publicized the creation of election integrity “war rooms” and rapid-response teams. These initiatives focused on visibility and coordination during election periods. Critics noted that underlying incentive structures and amplification mechanics remained largely unchanged. The measures reassured stakeholders without imposing durable constraints on system behavior. Governance was staged as readiness rather than structural reform.
Source: New York Times; internal Facebook disclosures

5. Universities — AI ethics boards without veto power

System: Institutional AI ethics committees
Date: 2021–2024

Many universities established AI ethics boards to oversee research and deployment. These bodies often lacked authority to block projects, enforce standards, or mandate changes. Recommendations were advisory, not binding. Institutions cited the existence of ethics review as evidence of responsible governance. In practice, the boards functioned as signaling mechanisms rather than control structures.
Source: Nature Machine Intelligence; Inside Higher Ed reporting


Failure State 7 — Responsibility Without Authority

When accountability is pushed onto actors who lack the power to alter system behavior.

1. OpenAI / ChatGPT — Fabricated case citations in Mata v. Avianca

System: ChatGPT
Date:
06/23

In Mata v. Avianca, a U.S. attorney submitted a legal brief containing multiple fabricated case citations generated by ChatGPT. When questioned by the court, the lawyer admitted relying on the system without verifying the sources. The judge sanctioned the attorney, emphasizing professional responsibility and diligence. The AI system itself provided no reliable indication that the cases were fictitious, nor any built-in verification affordance. Accountability was enforced entirely downstream, despite the lawyer having no authority over the system’s design or safeguards.
Source: U.S. District Court, S.D.N.Y.; New York Times, June 2023

2. Canadian Courts — AI-generated fake citations in filings

System: ChatGPT and similar LLM tools
Date: 07/23

Canadian courts encountered several instances where lawyers cited non-existent cases generated by AI tools. Judges reprimanded counsel and, in some cases, imposed cost penalties. Courts framed the incidents as failures of professional judgment rather than systemic tooling issues. Lawyers were expected to detect fabrication despite AI outputs appearing fluent and authoritative. Responsibility was assigned to individuals who lacked meaningful control over the systems producing the errors.
Source: Canadian Lawyer Magazine; Global News, July 2023

3. Air Canada — Attempted deflection of chatbot accountability

System: Air Canada website chatbot
Date: 02/24

After a tribunal ruled that Air Canada must honor incorrect refund information provided by its chatbot, the airline argued that customers should not rely on chatbot outputs. This position attempted to assign responsibility to users while retaining control over deployment and interface design. The tribunal rejected the argument, holding the company accountable for its digital agent. The case illustrates a common governance instinct: disclaim responsibility without relinquishing authority.
Source: Canadian Civil Resolution Tribunal decision; Business Standard, February 2024

4. Samsung — Internal data leaks followed by employee AI bans

System: ChatGPT and other public LLM tools
Date: 05/23

Employees at Samsung inadvertently shared sensitive internal data with public AI tools while seeking coding assistance. In response, Samsung restricted or banned employee use of generative AI tools. The governance response placed responsibility on individual workers to avoid misuse, rather than addressing system-level integration or safe-use design. Authority over tooling decisions remained centralized, while accountability was pushed onto employees. The incident exposed a gap between responsibility and control.
Source: TechCrunch, May 2023

5. Healthcare Systems — Clinicians warned not to rely on AI alerts

System: Clinical decision support AI tools
Date: 2022–2024

Hospitals increasingly deployed AI-driven alerts within electronic health record systems to flag patient risks. Clinicians were advised that these alerts were advisory only and that ultimate responsibility rested with them. At the same time, clinicians had little influence over alert thresholds, training data, or update cycles. Missed or incorrect alerts exposed providers to liability without granting authority to modify the system. Responsibility was imposed without corresponding control.
Source: STAT News; Journal of the American Medical Informatics Association