Abstract

As AI systems move from experimental tools to institutional infrastructure, governance has shifted from abstract ethics toward operational control. This paper surveys current academic and policy-oriented research on AI governance, with particular attention to risk management frameworks, human oversight requirements, and emerging architectural approaches for agentic systems. It argues that despite meaningful progress, most enterprise AI governance regimes fail to address a dominant failure mode: the informal transfer of authority under conditions of success. Against this backdrop, the Agora Commonplace Protocol (ACP) is introduced as a governance discipline focused on authority clarity, refusal legitimacy, and interruption rather than output optimization. ACP is not proposed as a replacement for existing frameworks such as the NIST AI Risk Management Framework or the EU AI Act, but as a complementary supervisory layer designed to preserve human responsibility where enterprise AI systems systematically erode it.


1. The current state of AI governance research

Over the past five years, AI governance research has consolidated around several shared concerns: risk identification, documentation, human oversight, and lifecycle accountability. What was once framed primarily as “AI ethics” is now increasingly treated as an engineering and institutional design problem.

1.1 Risk management as the dominant frame

In the United States, the National Institute of Standards and Technology’s AI Risk Management Framework (AI RMF) has become a central reference point for enterprise and government deployments. The framework emphasizes mapping, measuring, managing, and governing risks across the AI lifecycle, and its Generative AI Profile (NIST AI 600-1) explicitly extends these principles to large language models and foundation models.

Academic work aligned with this approach treats governance as a function of:

  • risk categorization,
  • mitigation controls,
  • documentation (model cards, data cards, factsheets),
  • and post-deployment monitoring.

This work has meaningfully improved transparency and consistency, but it assumes that identifying and mitigating risks is sufficient to constrain misuse.

The EU AI Act marks the most explicit legal codification of human oversight requirements, especially for “high-risk” systems. Article 14 mandates that systems be designed so humans can understand, supervise, and override them.

Academic responses to this requirement have raised a persistent concern: oversight is often nominal. Research on automation bias, decision fatigue, and time pressure shows that human operators frequently defer to system outputs even when override mechanisms exist. Oversight, in practice, becomes symbolic rather than substantive.

This gap between formal oversight and lived behavior is widely acknowledged, but rarely resolved.

1.3 Documentation, red-teaming, and evaluation

A third cluster of research focuses on documentation and stress testing. Model cards, dataset documentation, and large-scale red-teaming exercises—such as those coordinated around DEF CON—are framed as governance infrastructure.

These practices improve visibility into system behavior, but they largely operate after a system is already authorized to act. They do not reliably prevent outputs from being treated as de facto decisions in institutional contexts.


2. The architectural turn: governance as a control plane

A newer strand of research, emerging strongly in 2024–2026, addresses governance for agentic and tool-using AI systems. These systems perform multi-step actions, call external tools, and adapt dynamically to context.

In this literature, governance is increasingly described as a control plane:

  • policy enforcement layers,
  • authorization gates,
  • telemetry and provenance tracking,
  • interruption and rollback mechanisms.

This work recognizes that static policy documents and post hoc audits cannot govern systems that act in real time. Governance must be infrastructural.

This turn brings the field closer to ACP’s concerns, but it still largely treats governance as a mechanism for controlling system behavior, not for constraining how humans treat system outputs.


3. What remains missing: success as a failure condition

Across these bodies of work, a critical failure mode remains under-theorized: governance breakdown under success.

Most frameworks are triggered by error, harm, or incident. They assume failure is the signal that governance must intervene. Yet in real institutions, the most dangerous drift occurs when systems work well enough:

  • outputs are coherent,
  • time is saved,
  • friction is reduced,
  • and informal reliance expands.

Over time, language model outputs are cited, reused, and treated as authoritative without explicit authorization. Responsibility shifts quietly from humans to systems, even though no policy change occurred.

This phenomenon—sometimes described obliquely in the literature as overreliance or automation bias—is not well addressed by risk catalogs, documentation, or red-teaming. It is an authority problem, not an accuracy problem.


4. How ACP differs

ACP begins from a different premise: the primary governance risk of AI systems is not that they produce incorrect outputs, but that they restructure authority invisibly.

4.1 Authority-first rather than risk-first

Where enterprise frameworks ask “what could go wrong?”, ACP asks:

  • Who is authorized to ratify this output?
  • Who bears responsibility if it is acted upon?
  • Who can interrupt or refuse use?
  • Does refusal materially change system behavior?

This aligns ACP more closely with administrative law and institutional control theory than with ethics or safety checklists.

4.2 Helpfulness treated as a hazard class

Most AI governance assumes helpfulness is unambiguously good. ACP treats fluent, confident language as a risk vector, because it enables responsibility laundering: “the model said so.”

As a result, ACP explicitly legitimizes:

  • refusal,
  • non-engagement,
  • scope narrowing,
  • and silence.

These are treated as successful outcomes when engagement would increase dependency or normalize authority transfer.

4.3 Interaction as governance surface

ACP does not reside inside the model (as constitutions do), nor solely outside it (as policies and audits do). It operates at the interaction layer, shaping what kinds of outputs are allowed to exist in the first place.

An ACP-compliant artifact is not merely accurate; it is bounded, disclaimable, and resistant to reuse as doctrine.


5. Why ACP outputs can be materially better (under specific criteria)

“Better” here does not mean faster, more creative, or more accurate.

ACP outputs are better when the evaluation criteria include:

  • inspection survivability,
  • resistance to informal precedent formation,
  • clarity of human responsibility,
  • and reversibility under uncertainty.

In the exercises described earlier, ACP-governed behavior exhibited:

  • cross-domain continuity (essays, captions, test specs, meta-analysis),
  • consistent audience awareness,
  • explicit modeling of downstream misuse,
  • and willingness to suppress capability after success.

This pattern is not typical of enterprise chatbots, which optimize locally and reset context across tasks.


6. Coexistence, not replacement

ACP is not a competing “constitution” in the sense popularized by Anthropic. Constitutions aim to steer model behavior toward normative principles. ACP aims to constrain institutional use, regardless of model internals.

In practice, ACP could coexist with:

  • NIST AI RMF as a supervisory discipline,
  • EU AI Act compliance structures,
  • agentic AI control planes,
  • and constitutional or alignment-based models.

Its role is not to make models safer in isolation, but to make institutions safer in how they rely on models.


7. Research directions

If ACP is to be treated seriously as a governance approach, several research directions follow naturally:

  1. Refusal effectiveness studies
    Measure whether refusal and non-engagement reduce misuse more effectively than warnings or disclaimers.
  2. Authority clarity protocols
    Develop auditable checks for authority and interruption similar to security controls.
  3. Success-induced drift metrics
    Track how often successful outputs lead to scope creep, reuse, or informal normalization.
  4. Interface-level experiments
    Test whether interaction-layer constraints reduce overreliance more effectively than policy training.
  5. Comparative institutional trials
    Deploy ACP overlays in limited organizational settings and compare outcomes against standard enterprise AI governance.

Conclusion

The current AI governance literature has made substantial progress in risk management, documentation, and oversight. Yet it largely assumes that governance failures arise from error or malice. ACP highlights a different and equally dangerous dynamic: the quiet transfer of authority under success.

By treating refusal, constraint, and withdrawal as first-class outcomes, ACP reframes governance not as a promise of safety, but as a discipline of responsibility. Its value lies not in replacing existing frameworks, but in addressing what they systematically miss.

If AI systems are to become durable institutional infrastructure, governance must operate not only at the level of models and policies, but at the level of who is allowed to decide—and who is not.