In a 2025 paper, “Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming,” Anthropic’s red team investigates whether large language models can be made more robust against so-called “universal jailbreaks” — prompt strategies that systematically bypass safety mechanisms across many interactions. The paper introduces constitutional classifiers, safeguards trained on synthetic data generated from natural-language rules specifying permitted and restricted content, and reports results from thousands of hours of human red teaming and large-scale automated evaluation. Anthropic frames this work as evidence that model-level defenses against jailbreaks can be materially improved while remaining deployable at scale.
What question is Anthropic actually asking?
Anthropic is asking a technical containment question:
How do we prevent models from being induced, via adversarial prompting, to produce disallowed content at scale?
That is a coherent and non-trivial question. It is framed correctly for the object they have chosen: the model boundary.
But from an ACP perspective, this is already a narrowing move. The question presumes that the primary governance problem is model misbehavior under prompt pressure, rather than institutional reliance on model outputs. In other words, Anthropic is asking how to make the system say “no” more reliably — not how to ensure that when the system says anything at all, the authority to rely on it is legitimate, contestable, and owned.
So: the question is not wrong. It is incomplete in a predictable way.
Is Anthropic performing the right task?
Yes, they are performing the task they have set themselves extremely competently.
Training classifier layers against constitution-derived categories, stress-testing them with red-teaming, measuring overrefusal, compute overhead, and jailbreak success rates — all of that is solid engineering work. From a pure ML-systems perspective, this is careful, honest, and technically serious.
From an ACP perspective, however, the task itself is orthogonal to the core failure mode.
Anthropic is optimizing content compliance at the interface. ACP is concerned with authority allocation downstream. A system can be perfectly jailbreak-resistant and still be institutionally catastrophic if its outputs are treated as self-authorizing, unappealable, or responsibility-free.
This is the key misalignment:
- Anthropic treats refusal as a property of the model-plus-classifier stack.
- ACP treats refusal as a governance act that must be owned, contestable, and reviewable.
The moment a classifier’s refusal becomes “the answer,” ACP would flag a silent authority transfer — regardless of how accurate or robust the classifier is.
So: they are performing the task they believe matters.
ACP would say it is not the task that decides institutional safety.
Is the process and are the results effective (in ACP terms)?
Technically: yes, within scope.
Institutionally: only weakly, and with new risks introduced.
The paper is unusually candid in ways ACP would respect: they report overrefusal, compute costs, and the fact that a universal jailbreak was eventually found. That honesty matters. It signals that the authors themselves do not believe they have “solved alignment.”
But from an ACP standpoint, the results demonstrate something different from what the paper foregrounds.
They demonstrate that:
- robustness gains are real but temporary,
- adversarial pressure reasserts itself,
- and every new enforcement layer becomes a new surface of authority.
The system becomes harder to break, but also harder to question. A refusal produced by a classifier trained on synthetic constitutional data is not easily contestable by a user, a reviewer, or an institution. There is no standing, no appeal path, no reason-giving obligation beyond “the classifier flagged it.”
From ACP’s view, that is a governance regression disguised as a safety improvement.
So while the process is effective at reducing a particular class of failures (prompt-induced disallowed outputs), it simultaneously amplifies the risk of authority laundering by making refusals feel neutral, technical, and final.
What does ACP conclude — and what would it do differently?
ACP’s conclusion is not that Anthropic is misguided or acting in bad faith. It is that Anthropic is operating inside the dominant “Big AI” frame, where governance is treated as something you add to the model boundary rather than something you impose on institutional use.
From an ACP perspective:
- Anthropic is asking how to stop the model from saying the wrong thing.
- ACP asks how to stop institutions from letting the model decide what matters.
Those are different problems.
ACP would not reject Anthropic’s work. It would demote it. Classifiers become tooling — useful, fallible, explicitly non-authoritative — embedded in a larger system where:
- refusals are explainable and appealable,
- reliance is explicitly authorized by humans,
- and jailbreak resistance is treated as risk reduction, not legitimacy.
The most ACP-aligned reading of the paper is therefore this:
It is evidence that technical constraint helps but does not govern; that enforcement without institutional process creates new authority hazards; and that robustness does not eliminate the need for refusal as a human-owned act.
In that sense, the paper is actually strong supporting evidence for ACP’s core claim, even though it does not articulate that claim itself.
Bottom line
Anthropic is asking a reasonable but partial question, performing the task it implies with high competence, and producing results that are real but insufficient for governance.
From an ACP perspective, the work is best understood not as an answer to alignment, but as a case study in why alignment cannot be solved at the model boundary alone.
The danger is not that Anthropic’s approach fails. The danger is that institutions treat its partial success as permission to stop asking harder questions. ACP exists to keep those questions alive.
Member discussion: