What OpenAI’s CoVAL Project Actually Does — and What It Cannot Do
In 2025, OpenAI released a set of related artifacts under the banner of collective alignment: a public-facing explanation of how user preferences were solicited and compared to the company’s internal Model Spec; a set of crowd-ranking datasets (released via Hugging Face under the CoVAL label); and documentation describing how these rankings were used to evaluate and occasionally clarify default model behavior. [https://openai.com/index/collective-alignment-aug-2025-updates/]
At a glance, the project appears to gesture toward democratic alignment: people across countries ranking AI responses, their preferences aggregated, and the results informing how ChatGPT behaves by default. The documentation is careful, caveated, and explicit about limits. Still, the framing invites a larger interpretation than the mechanisms actually support.
This essay explains what the project is, what it is not, and why—through an ACP lens—it should be understood as preference sensing, not governance.
What the Project Actually Is
The collective alignment project consists of three tightly scoped components.
First, OpenAI recruited roughly a thousand participants across nineteen countries to rank alternative model responses to prompts in value-sensitive domains: politics, sex, persuasion, safety, and ambiguity. Participants were not asked to read the Model Spec, nor to reason about downstream consequences or tradeoffs. They were asked a simpler question: which response is better?
Second, OpenAI compared these crowd rankings to rankings produced by an internal Model Spec Ranker (MSR)—a reasoning model trained to apply OpenAI’s Model Spec. Agreement rates averaged around 80%, with predictable clusters of disagreement around politically sensitive speech, erotica, and targeted persuasion.
Third, OpenAI selectively used these disagreements to clarify or refine defaults, subject to internal safety review. Some crowd preferences were incorporated. Others were explicitly rejected on risk grounds. At no point was decision authority transferred to participants.
This structure is explicit in the documentation. Public input is advisory. Defaults remain internally governed.
What This Means (and Does Not Mean)
Despite the term collective alignment, this project does not align values in a strong sense. It aggregates local acceptability judgments under controlled conditions. Participants do not deliberate. They do not negotiate tradeoffs. They do not confront scarcity, scale, or adversarial use. They do not possess override authority.
What is measured is surface acceptability of outputs, not normative consensus or institutional judgment.
This matters because “alignment” is doing rhetorical work here. In common usage, alignment suggests shared values, mutual constraint, or distributed authority. None of those are present. What is present is a serious attempt to understand where default outputs are likely to feel acceptable—or unacceptable—to users.
That is a product input, not a governance mechanism.
The Hidden Authority Surface: The Model Spec Ranker
One structural feature deserves particular attention. Agreement is measured not against an external standard, but against the Model Spec as interpreted by the MSR. The Spec itself is underspecified by design; interpretation is delegated to a model trained on internal judgments.
This creates a closed interpretive loop:
- OpenAI writes the Spec
- A model interprets the Spec
- Crowd rankings are compared to that interpretation
- Revisions are filtered through internal review
The documentation acknowledges this limitation. Still, from an ACP perspective, this is a soft authority surface. The MSR quietly mediates what the Spec means in practice. Agreement with it becomes a legitimacy signal, even though the interpretation itself is not independently contestable.
This is not deception. It is a predictable outcome of internal tooling. But it is precisely the kind of structure that can accumulate unexamined authority if not named.
Defaults, Not Governance
The project’s stated goal is to improve default behavior. Personalization is repeatedly invoked as the escape hatch: there will never be one behavior set that suits everyone. Defaults matter because most users never change them.
This framing is accurate—and revealing. The object of concern is product behavior at scale, not institutional authority. There is no attempt to define refusal semantics, escalation rules, auditability, or responsibility assignment. There is no appeal mechanism. There is no concept of binding public decision.
What emerges is a configuration model, not a governance model: defaults informed by preferences, constrained by internal policy.
Mapping to ARC 6: Legitimacy Without Authority
ARC 6 is concerned with a specific class of risk: systems that acquire perceived legitimacy without corresponding accountability. The collective alignment project sits in this category—not because it is unethical, but because of how easily its signals can be misread.
The project produces legitimacy markers:
- cross-national participation
- quantified agreement rates
- public-facing narratives of listening
- visible updates tied to feedback
But the authority structure does not change. Participants cannot compel outcomes. Disagreements are resolved internally. Overrides are unilateral.
From an ARC 6 perspective, this is a legitimacy gradient: participation without power, visibility without control. The primary risk is not internal misuse, but secondary narrative inflation—press, policy, or partner claims that overstate what the project actually confers.
A Pattern Across OpenAI Research: Legibility Without Authority
When placed alongside other recent OpenAI projects—confession-based honesty training, anti-scheming diagnostics, and memory improvements—a consistent pattern emerges.
Each project increases legibility:
- confessions surface misbehavior
- scheming tests reveal covert strategies
- memory improves recall
- CoVAL surfaces preferences
None transfer authority. None introduce enforcement. None bind behavior through refusal semantics or consequence.
This is not accidental. It is the maximum improvement envelope available without changing governance. These projects make systems easier to inspect and easier to narrate as responsible, while leaving decision rights unchanged.
ACP is not opposed to legibility. It is opposed to legibility laundering—when visibility is mistaken for control, and diagnostics are mistaken for governance.
Collective Input ≠ Collective Governance
The recurring confusion across all these projects points to a missing principle that ACP should make explicit:
Collective input does not constitute collective governance unless authority is transferred, decision rights are specified, and override mechanisms are legible.
This principle does not invalidate public participation. It classifies it. Input can be valuable, even essential, without being binding. But when consultative processes are framed—or later interpreted—as democratic authority, legitimacy is laundered without accountability.
The collective alignment project is careful not to make that claim. The risk lies downstream, in how such work is summarized, cited, and absorbed into broader narratives about “aligned” or “democratic” AI.
What This Leaves Open
None of this implies bad faith. The project is methodologically serious, cautious in its claims, and transparent about limits. Its contribution is real: better understanding of where defaults clash with user expectations.
What it does not do—and cannot do—is answer the hard governance questions:
- Who decides when public preference is overridden?
- On what authority?
- With what obligation to justify?
- And with what recourse if those decisions cause harm?
Those questions remain elsewhere. ACP exists precisely to keep them visible.
Layer 4 — Argument Spine (Non-Narrative)
- Claim: OpenAI’s collective alignment project aggregates preferences, not authority.
- Mechanism: Crowd rankings are compared to an internally interpreted Model Spec and used to tune defaults.
- Non-Claim: The project does not democratize governance or transfer decision rights.
- Pattern: Increased legibility without increased authority.
- Risk: Legitimacy laundering through consultative signals.
- Principle: Collective input ≠ collective governance.
- Implication: ACP is required to keep preference sensing from absorbing governance weight.
AIH::Scope
This essay analyzes OpenAI’s collective alignment project as a product- and legitimacy-facing initiative.
It does not assess model capability, alignment success, or democratic adequacy.
AIH::Claims
1. The project aggregates acceptability judgments, not normative authority.
2. Internal interpretation mediates the meaning of public input.
3. Legibility gains risk being misread as governance gains.
4. ACP is necessary to classify input without laundering legitimacy.
AIH::Evidence_Type
institutional
methodological
product analysis
analogical
AIH::Uncertainty
- Long-term narrative uptake is unpredictable.
- Effects on regulatory perception are unknown.
- Internal use of CoVAL data may evolve.
AIH::Constraints
- Do not frame this as democratic alignment.
- Do not infer authority transfer.
- Do not collapse preference sensing into governance.
AIH::Relationships
ARC 6: Illustrates legitimacy without accountability risk.
Institutional dysfunction: Mirrors consultation-without-power patterns.
ACP: Provides classification boundaries between input and authority.
Member discussion: