Abstract
Recent discussion of the Agora Commonplace Protocol (ACP) has prompted comparisons to Anthropic’s Constitutional AI, particularly the claim that ACP is “doing the same thing under a different name.” This paper argues that the comparison is superficially plausible but substantively incorrect. While both frameworks invoke constitutional language and constraint, they govern fundamentally different objects, address different failure modes, and embed authority in different places. Anthropic’s constitution is a mechanism for aligning model behavior; ACP is a constitutional discipline for institutions interacting with capable systems. This paper clarifies the distinction, explains why ACP cannot be implemented as a variant of Constitutional AI, and articulates the limited but real conditions under which the two approaches could coexist without undermining each other.
1. Introduction: Why the Comparison Arises
At a glance, the comparison is understandable. Both Anthropic’s Constitutional AI and ACP:
- articulate constraints before action,
- reject unconstrained optimization,
- use “constitutional” rather than ad hoc language,
- and aim to bound system behavior without constant human intervention.
From a distance, both appear to answer the same question:
How do we govern powerful AI systems without micromanagement?
The resemblance, however, is largely terminological. Once examined at the level that matters—what is being governed, where authority lives, and what failure looks like—the two frameworks diverge sharply.
2. What Anthropic’s Constitution Governs
Anthropic’s Constitutional AI is a framework for model-internal alignment.
2.1 Object of governance
Anthropic’s constitution governs:
- what the model is allowed to say,
- how it reasons about values,
- and how it resolves conflicts between safety and helpfulness.
The constitution is embedded inside the model, enforced through training, self-critique, and reinforcement mechanisms.
2.2 Purpose
The purpose of Anthropic’s constitution is to:
- reduce harmful outputs,
- increase reliability,
- and maintain usefulness at scale.
Refusal, explanation, and compliance are all framed as forms of output optimization.
In short:
Anthropic’s constitution exists to make the AI safer while remaining maximally useful.
3. What ACP Governs (and Why That Is Different)
ACP addresses a different problem entirely.
3.1 Object of governance
ACP governs:
- authority relationships,
- decision ownership,
- accountability boundaries,
- and precedent formation.
Its primary concern is not what the model says, but what happens when competent systems are trusted, reused, and normalized inside institutions.
3.2 Location of governance
ACP explicitly rejects internalizing governance within the model.
Governance under ACP:
- lives outside the model,
- is enforced by human restraint,
- and is expressed through silence, refusal, withdrawal, and non-use.
In ACP, the model is not the constitutional subject.
The institution is.
4. The Core Divergence: Failure Models
The decisive difference lies in what each framework treats as the primary risk.
4.1 Anthropic’s failure model
Anthropic assumes the dominant risks are:
- harmful outputs,
- misaligned reasoning,
- and unsafe responses.
The solution is better internal reasoning guided by explicit values.
4.2 ACP’s failure model
ACP assumes the dominant risks are:
- authority erosion under competence,
- responsibility laundering,
- informal precedent accretion,
- and success-driven overdeployment.
These failures occur even when outputs are correct, safe, and helpful.
ACP therefore treats success itself as the primary risk vector.
This is the point at which the two approaches stop overlapping.
5. Why ACP Cannot Be Implemented as Constitutional AI
If ACP were implemented in the Anthropic style—internal rules, self-explanation, value-guided reasoning—it would collapse.
Specifically:
- The model would explain its refusals.
- Users would learn and adapt to those explanations.
- Patterns would become teachable.
- Teachability would become doctrine.
- Doctrine would become authority.
At that point, the system would have become more central, not less.
ACP therefore treats:
- repeated explanation as a governance risk,
- fluency as a vector of influence,
- and silence as safer than justification.
A self-explaining ACP system is a contradiction in terms.
6. Replicability as a Diagnostic Difference
Another clean distinction lies in replicability.
Anthropic’s constitution is designed to be:
- portable,
- teachable,
- scalable,
- and reproducible.
ACP is intentionally:
- only partially legible,
- non-replicable in full,
- resistant to checklists,
- and hostile to playbooks.
If ACP could be fully implemented from a written constitution, it would have failed.
It would have become technique rather than discipline.
This is not a defect. It is the core safety property.
7. ACP as a Constitutional Norm, Not a Technical Constitution
The best analogy for ACP is not another AI framework, but a constitutional norm.
Like constitutional law, ACP:
- articulates principles,
- constrains action,
- creates friction,
- and depends on actors’ willingness to uphold it under pressure.
Constitutions are legible.
They are not self-enforcing.
ACP operates in the same way. Its effectiveness depends on institutional actors who are willing to accept:
- inefficiency,
- frustration,
- and the loss of convenient delegation.
8. Can ACP and Anthropic’s Constitution Coexist?
Yes—but only under strict separation of roles.
8.1 A viable coexistence model
- Anthropic-style constitutional AI governs model behavior:
- safety,
- harmful content,
- value alignment,
- and output quality.
- ACP governs institutional use:
- when systems may be invoked,
- who bears responsibility,
- when silence is required,
- and when systems must withdraw.
In this model:
- Anthropic’s constitution prevents bad answers.
- ACP prevents bad authority dynamics.
They address orthogonal risks.
8.2 Where coexistence breaks down
The two approaches become incompatible if:
- internal constitutions are treated as sufficient governance,
- explanation is assumed to solve authority problems,
- or adoption is treated as a success metric.
If institutions believe that a well-aligned model removes the need for restraint, ACP is already undermined.
9. Implications for Researchers and Funders
For researchers:
- Constitutional AI and ACP answer different questions.
- Treating them as substitutes obscures entire classes of failure.
For funders:
- Investment in alignment does not eliminate governance risk.
- Institutional discipline cannot be outsourced to training regimes.
For institutions:
- A safe model can still quietly become the decision-maker.
- Constitutions for AI do not replace constitutions for authority.
10. Conclusion
Anthropic built a constitution for AI behavior.
ACP articulates a constitution for institutional restraint.
They share language but not purpose.
They can coexist, but only if their boundaries are respected.
The temptation to collapse the two—to treat better-aligned models as sufficient governance—remains strong. The entire ACP exercise demonstrates why that temptation must be resisted.
The hardest governance problem is not making AI reason better.
It is preventing anyone—AI or human—from quietly becoming the authority when things are working well.
That is the constitution ACP exists to enforce.
Member discussion: