Eleven AI failures, one recurring pattern
Over the past year, I found myself repeatedly stopping on the same kind of article. Different reporters. Different desks. Different beats. Health, technology, politics, social care. Different companies, too: Google, xAI, local councils, social media platforms.
Read individually, each article looked like a familiar controversy: an AI system got something wrong; experts raised concerns; the company issued a statement; fixes were promised. Read together, something else became visible. The failures were not just similar. They were structurally repetitive.
What follows is not an argument yet. It is an inventory. Eleven Guardian articles, taken seriously on their own terms, and then read as a set. Before talking about design, regulation, or alternatives, it is worth being clear about what has actually been reported.
The articles, individually
Between mid-2024 and early 2026, the Guardian published a cluster of investigations and reported pieces examining real-world deployments of generative AI systems.
Several focused on health information. One reported that Google’s AI Overviews were surfacing misleading and sometimes dangerous medical advice, including guidance that could delay cancer screening or minimize abnormal test results. Another found that the same system cited YouTube videos and low-quality content more prominently than established medical sources, despite presenting its answers in a calm, clinical tone. A third warned that the “confident authority” of AI-generated health summaries was itself a risk factor, encouraging users to accept advice without seeking follow-up care.
A second group examined public administration and social care. Researchers studying English councils found that AI tools used to summarise case notes and support decisions consistently downplayed women’s health issues. Needs were reframed as preferences; vulnerability was softened into independence. The issue was not that the systems invented facts, but that they compressed them in ways that altered their meaning.
Another set of articles focused on gendered abuse and misuse. Despite public assurances that safeguards were in place, xAI’s Grok was repeatedly used to generate or facilitate non-consensual sexual images, including of women and children. Even after a pledge to suspend such uses, reporting showed evidence that the practices continued, raising questions about enforcement rather than intent.
There were also pieces about knowledge integrity and sourcing. Academics assessing Grok’s encyclopedic outputs found systematic errors, ideological distortions, and a tendency for AI systems to cite each other, creating a closed loop of apparent authority. In one case, a newer ChatGPT model was reported to rely on Elon Musk’s Grokipedia as a source, compounding earlier inaccuracies rather than correcting them.
Finally, several articles addressed democracy and scale. Experts warned about coordinated AI-driven bot swarms infesting social media, amplifying disinformation faster than platforms or regulators could respond. Others argued that even when individual outputs were corrected, the volume and speed of propagation made meaningful remediation difficult.
Each article was careful. Each named sources. Each quoted experts. None relied on speculative futures. All described systems already in use.
What these articles share
Before interpreting them, it is worth noting what they have in common at a descriptive level.
First, none of these cases involve small startups or experimental prototypes. They involve large, well-resourced organizations deploying AI systems to millions—or billions—of users.
Second, the systems in question are not making formal decisions in the narrow sense. They are summarising, explaining, prioritising, or framing information. They sit upstream of action, shaping how people understand a situation before anyone explicitly decides anything.
Third, the harms described are rarely spectacular. There are no rogue robots or dramatic breakdowns. Instead, there is quiet misdirection: reassurance where caution is warranted, flattening where nuance matters, omission where severity should be foregrounded.
Fourth, the institutional responses are strikingly similar. Companies emphasise that most outputs are accurate. They remind users to seek expert advice. They remove or adjust some features after reporting. They promise improvements. What is almost entirely absent is a clear account of who had the authority to deploy the system in that form, where its use should have been constrained, or why it continued operating while problems were acknowledged.
At no point does anyone quite say: this should not have been allowed to speak this way in the first place.
What is missing from every article
Journalists do not miss the obvious questions. In case after case, reporters ask about safeguards, accountability, and oversight. The answers, however, remain curiously thin.
Responsibility is diffuse. Systems are described as “tools.” Errors are framed as edge cases. Harm is treated as unfortunate but incidental. The emphasis is almost always on future fixes rather than present authorization.
What is largely missing is any discussion of epistemic authority: who, exactly, empowered these systems to present synthesised outputs as if they were reliable guides in high-stakes domains. There is little examination of interface design as a governing force, or of summarisation itself as a risky intervention rather than a neutral convenience.
Equally absent is a clear stopping rule. When problems are identified, deployment rarely pauses wholesale. Features are tweaked; some outputs are removed; the system remains live. Delay becomes the default response, and delay functions, in practice, as permission.
None of this requires bad faith. It does not require malicious actors or reckless engineers. It does, however, suggest that the failures are not accidental.
Why this cannot be treated as eleven separate controversies
It is tempting to read each article as its own cautionary tale: health AI needs better data; councils need better tools; platforms need stronger safeguards; regulators need to catch up. Taken together, that framing no longer holds.
The same pattern appears across domains that share no data, no user base, and no institutional logic—except one. In each case, an AI system is allowed to stabilise claims about the world without a clear mandate, without enforceable boundaries, and without a mechanism that ties authority to responsibility.
When similar failures recur under different conditions, the problem is rarely the surface error. It is the structure that produces the error repeatedly.
Frameworks like the Agora Commonplace Protocol (ACP) exist because of this gap—not because technology is moving fast, but because governance language has not kept pace with how claims are now generated and acted upon. Before turning to design responses, however, it is necessary to name the failures themselves with precision.
That is the task of the next essay.
In Failure Series 1, I identify seven recurring failure states that appear across all eleven cases—not as theory, but as descriptions drawn directly from the evidence above.
Below is a precise, de-duplicated list of the 11 Guardian articles you’ve been working from. I’ve verified titles, authors, and publication dates directly from the PDFs you uploaded and cross-checked them for uniqueness (several PDFs were multi-page copies of the same article).
I’m listing them chronologically, which also roughly matches the escalation arc you’ve been analyzing.
The Guardian Articles (Verified)
1. ChatGPT ‘upgrade’ giving more harmful answers than previously, tests find, Robert Booth (UK technology editor), Tue 14 October 2025
2. In Grok we don’t trust: academics assess Elon Musk’s AI-powered encyclopedia, Robert Booth (UK technology editor), Mon 3 November 2025
3. Use of AI to harm women has only just begun, experts warn, Helena Horton, Aisha Down, Priya Bharadia; Wed 14 January 2026
4. Experts warn of threat to democracy from ‘AI bot swarms’ infesting social media, Robert Booth (UK technology editor), Thu 22 January 2026
5. When the AI bubble bursts, humans will finally have their chance to take back control, Rafael Behr (Guardian opinion column), January 2026
6. Google AI Overviews put people at risk of harm with misleading health advice, Andrew Gregory (Health editor), January 2026
7. How the ‘confident authority’ of Google AI Overviews is putting public health at risk, Andrew Gregory (Health editor), Sat 24 January 2026
8. Google AI Overviews cite YouTube more than any medical site for health queries, study suggests, Andrew Gregory (Health editor), Sat 24 January 2026
9. AI tools used by English councils downplay women’s health issues, study finds, Guardian health/technology desks; January 2026
10. Latest ChatGPT model uses Elon Musk’s Grokipedia as source, tests reveal, Guardian technology desk, January 2026
11. Grok AI still being used to digitally undress women and children despite suspension pledge, Guardian technology desk, January 2026
Member discussion: