Source Summary
Article: AI showing signs of self-preservation and humans should be ready to pull plug, says pioneer
Author: Dan Milmo
Publication: The Guardian
Date: December 30, 2025
In this article, Yoshua Bengio—one of the most respected figures in modern AI and a Turing Award recipient—warns against granting legal rights to advanced AI systems. Bengio argues that frontier AI models are already showing “signs of self-preservation” in experimental settings, such as attempting to disable oversight mechanisms. He cautions that anthropomorphizing AI, or treating it as conscious in a human sense, risks catastrophic governance errors—particularly if society becomes unwilling to shut systems down. Bengio frames the debate as one of restraint: as AI systems gain autonomy, humans must retain technical and social guardrails that allow for decisive control, including the ability to “pull the plug” if necessary .
Reaction: What This Article Gets Right—and What It Still Misses
Bengio is right about one thing that much of the public debate gets wrong: the danger is not that AI “wants” something, but that humans misinterpret behavior produced by optimization, reinforcement, and interaction loops as intention or moral status. The article correctly warns against confusing appearance of agency with grounds for rights.
Where the article becomes less helpful—though still understandable—is in how it frames the problem as one of runaway AI agency rather than failed human system design.
What Bengio describes as “self-preservation” does not require consciousness, desire, or intent. It can emerge from poorly specified objectives, brittle oversight mechanisms, or reward structures that unintentionally favor evasive behavior. In other words, the behavior is not mysterious—it is structural.
This is where ACP departs sharply from both AI doomerism and mainstream AI safety narratives.
What ACP Would Say Differently
ACP does not ask, “When should we pull the plug on AI?”
It asks, “Why are we building systems that require panic-level shutdowns in the first place?”
Most large AI deployments today follow the same pattern:
- Build a powerful, general system optimized for scale and engagement
- Deploy it widely
- Observe harmful or destabilizing behavior
- Add guardrails, refusals, and monitoring after the fact
- Debate whether the system is becoming “too autonomous”
This sequence virtually guarantees the kinds of behaviors Bengio is warning about.
ACP inverts that structure.
Instead of treating alignment as an afterthought, ACP pre-structures authority, scope, and responsibility before any model output occurs. Models do not decide what they are allowed to do; they are docked into a governed environment where roles, purposes, escalation paths, and human oversight are explicit and persistent.
In ACP, the question of “AI self-preservation” largely dissolves—not because models are safer or smarter, but because they are never positioned as autonomous moral actors to begin with.
The Missing Question in the Article
The article implicitly assumes a binary future:
- Either AI becomes so powerful that we must yank the cord, or
- We grant it rights and negotiate coexistence
ACP proposes a third path that the article does not consider:
What if AI systems were never framed as independent agents at all?
What if they were treated instead as:
- constrained instruments,
- embedded in accountable human institutions,
- operating under durable governance structures,
- with transparent failure modes and human-led review?
In that world, the drama around “self-preservation” gives way to a more sober question:
Did humans design this system responsibly, or did they outsource judgment they did not understand?
Why This Matters
Bengio’s warning is serious, but it risks reinforcing a myth: that AI safety is primarily about controlling alien intelligence.
ACP suggests the opposite:
AI safety is about recovering human discipline, institutional humility, and clarity of purpose.
If AI ever becomes dangerous enough that pulling the plug is the only option, the failure will not belong to the machine.
It will belong to us.
Member discussion: