1) Where the research has converged

Across disciplines, “AI governance” has stopped meaning only ethics principles and started meaning operational control surfaces: who can authorize use, what gets logged, how failures are detected, and how humans can actually intervene.

Three clusters show up repeatedly:

A. Risk-management frameworks becoming implementation targets (not just guidance).
NIST’s AI RMF has become a de facto reference point for U.S.-centric enterprise governance, and the Generative AI Profile (NIST AI 600-1) pushes the RMF into GenAI-specific risk categories and controls. (NIST)

B. “Human oversight” treated as both a legal requirement and an engineering problem.
The EU AI Act makes human oversight a formal obligation for high-risk systems, but a large part of the academic discussion is that oversight can be nominal, brittle, or undermined by automation bias and system complexity. (Artificial Intelligence Act EU)

C. Documentation, traceability, and evaluation as governance primitives.
Model/dataset documentation (Data Cards, model cards, “factsheets”) and structured red-teaming are being treated as governance infrastructure: you can’t manage what you can’t describe, and you can’t describe what you didn’t log. (ACM Digital Library)

2) The “agentic AI” shift: governance as a control plane

The 2025–2026 wave you’d want on your radar is governance for agentic systems (multi-step, tool-using, semi-autonomous workflows). Papers are increasingly explicit that classic AppSec + policy documents won’t scale to improvisational systems, so they propose governance layers that look like control planes: policy enforcement, telemetry, provenance, interruption, anomaly detection, and lifecycle assurance. (arXiv)

This is conceptually close to what you’ve been calling an “overlay”: governance decoupled from the model’s “helpfulness” incentives.

3) Constitutional approaches: “principles” inside the model vs governance outside it

Anthropic’s “Constitutional AI” popularized the idea of encoding a set of principles (a “constitution”) that steers model behavior. That line of thinking has also sparked broader work on legitimacy and public-facing constitutions. (arXiv)

But most of that work—by design—optimizes for model behavior (safer outputs), not for institutional accountability (who is authorized to treat outputs as decisions, how responsibility is prevented from drifting).

This distinction matters for ACP.


What’s still structurally missing (even in good governance research)

If you read the frameworks, tool proposals, and red-teaming literature together, three gaps recur:

Gap 1: Governance is often “additive,” not “interruptive”

A lot of governance is layered as reporting and review (documentation, evaluations, red teaming transparency reports), but it does not reliably produce a material change in system behavior when authority is absent or contested. The system still answers; humans may or may not notice the risk signal. (Humane Intelligence)

Gap 2: Oversight is required, but “meaningful oversight” is rarely enforced

Legal requirements and design guidance emphasize oversight and override capacity, but researchers also argue oversight can be infeasible as systems become more complex and autonomous—and humans over-trust outputs in practice. (ScienceDirect)

Gap 3: The biggest institutional failure mode is success

Many approaches assume failure is the trigger (incidents, model misbehavior, audit findings). But in enterprises, the dangerous drift often comes from repeated adequate performance: adoption expands informally, boundaries blur, and “the model said so” becomes a social fact. (This theme also shows up in public red-teaming discussions: scaling testing is valuable, but it can create false confidence if governance is treated as “we did the exercise.”) (Harvard Data Science Review)


How ACP differs (and why it can be materially better—under specific criteria)

To stay intellectually honest: “better” depends on what you’re optimizing for.

If the target metric is task throughput, classic enterprise AI plus documentation/red-teaming is usually “better.” If the target metric is inspection survivability and non-laundered responsibility, ACP can be better because it treats governance as a discipline of authority, not as a set of principles or post hoc evaluations.

Here are the crisp differentiators:

1) ACP is authority-first; most enterprise AI governance is risk-first

  • Risk-first: enumerate harms, add mitigations, document decisions. (NIST Technical Series)
  • Authority-first (ACP): before content quality, establish who can ratify claims, who can interrupt, and whether refusal changes behavior.

That’s closer to administrative control theory than to “responsible AI checklists.”

2) ACP treats “helpfulness” as a hazard class

A lot of governance assumes the model is a tool and problems arise when it errs. ACP treats the model’s strongest social affordance—coherent helpful language—as a primary vector for responsibility laundering and precedent accretion.

This aligns with (and extends) academic concerns about automation bias / overreliance, but ACP operationalizes it as: non-engagement thresholds, scope containment, and refusal legitimacy. (ScienceDirect)

3) ACP’s core output is not “an answer,” but an institutionally usable artifact with traceable boundaries

The research literature tends to split:

  • inside-model constraints (e.g., constitutions), or
  • external governance infrastructure (documentation, red-teaming, control planes). (arXiv)

ACP is trying to make the interaction layer itself a governance surface: “this output cannot be treated as policy,” “this requires named authority,” “here is where escalation is required,” etc.—and to do so in language that survives inspection.

That’s a different object than a safety policy, and different from a model card.


Concrete examples worth citing (for the paper’s “named institutions” requirement)

If you want the survey to feel non-generic, these are strong anchors:

  • NIST AI RMF + GenAI Profile (NIST AI 600-1) as the U.S. reference spine for enterprise implementation. (NIST)
  • EU AI Act human oversight requirements as the clearest “oversight is law” artifact (and a site of ongoing feasibility debate). (AI Act Service Desk)
  • DEF CON / public red-teaming transparency work as the canonical “scaled stress testing” example—and also a demonstration that testing doesn’t equal governance. (Humane Intelligence)
  • MIT AI Governance Mapping as an institutional attempt to index the governance landscape (useful for demonstrating the field’s breadth and fragmentation). (MIT AI Risk Repository)
  • Agentic governance control-plane proposals (e.g., Kubernetes-native governance, semantic telemetry, dynamic authorization) as the “architecture is shifting” evidence. (arXiv)

Research agenda: what to study next if ACP is taken seriously

If ACP is to be more than an idiosyncratic discipline, the research questions become testable:

  1. Refusal effectiveness as a governance control
    When does refusal (or constrained engagement) measurably reduce downstream misuse, escalation failures, or responsibility laundering—compared to “best effort answers + warnings”?
  2. Authority clarity protocols as measurable interventions
    Can we define and validate “authority clarity” checks the way we validate security controls (pass/fail, audit logs, repeatability)? This connects naturally to control-plane governance work for agents. (arXiv)
  3. Success-induced drift metrics
    Develop metrics for “hinge formation”: frequency of re-invocation after success, scope creep rates, reuse of phrasing as informal doctrine, and time-to-normalization in organizations.
  4. Interface-level governance experiments
    A/B test whether UI-enforced boundary disclosures (who decides, who bears risk, what counts as ratified) reduce overreliance more effectively than policy training. This is adjacent to the human oversight feasibility concerns in the literature. (ScienceDirect)
  5. Comparative institutional trials
    Run ACP-style overlays in a few real workstreams (legal review, procurement, HR, incident response) and compare:
  • rework rate,
  • escalation correctness,
  • audit defensibility,
  • and incidence of “model-as-authority” language in downstream artifacts.

Bottom line comparison (ACP vs “enterprise AI governance”)

  • Enterprise governance research is increasingly strong on risk catalogs, documentation, red-teaming, and agent control planes. (NIST Technical Series)
  • ACP is strongest where the enterprise stack is weakest: preventing informal authority transfer under success conditions, and forcing an explicit account of who can ratify and interrupt.