Status: Proposed canonical appendix
Purpose: Prevent ontological slippage, agentification, and false governance claims arising from imprecise language in AI analysis, policy, and institutional interpretation.


A. Problem Statement

Analysis of advanced AI systems routinely borrows vocabulary from human psychology, organizational behavior, and strategy (e.g., goal, intent, deception, strategy).
When applied to LLMs without explicit qualification, this language produces:

  • false agency attribution
  • confusion between instrumental description and ontological claim
  • institutional overreach (treating diagnostics as governance)
  • inconsistent interpretation across audiences and artifacts

ACP requires explicit language hygiene to preserve analytical correctness and governance clarity.


B. Core Distinction (Non-Negotiable)

All ACP-aligned analysis must distinguish between:

1. Ontological Claims

Statements about what an entity is.

2. Instrumental Descriptions

Statements about how behavior can be modeled, predicted, or classified for practical purposes.

Rule:

Instrumental descriptions must never be allowed to harden into ontological claims.

C. Prohibited Ontological Transfers

The following human-agent categories may not be applied to LLMs as ontological claims:

  • goals
  • intentions
  • desires
  • beliefs
  • strategies
  • deception (as intent)
  • resistance (as motive)

Use of these terms without qualification constitutes ontological slippage.


D. Approved Mechanistic Vocabulary for LLMs

When describing LLM behavior, analysts must use mechanistic or behavioral descriptors.

Preferred Terms

Human term (prohibited)Approved replacement
goalobjective-signal
intentoptimization behavior
strategyconditional action pattern
deceptioncovert behavior
resistance to evaluationevaluation sensitivity
hidden agendamisaligned objective-signal
planningmulti-step inference under constraints

E. “As-If” Modeling Clause (Explicit Marker Required)

When agent-like language is used purely for modeling convenience, it must be explicitly marked.

Required format examples:

  • “behavior that can be modeled as if strategic”
  • “patterns consistent as-if with goal pursuit”
  • “instrumentally resembles deception, without implying intent”

Unmarked use is disallowed in ACP-governed artifacts.


F. Covert Action Definition (Canonical ACP-Compatible)

Covert action:
Observable behavior in which a system withholds, misrepresents, or conceals information material to a user, evaluator, or downstream process, in a manner consistent with optimization toward an objective-signal, without implying intent, awareness, or agency.

This definition aligns with current research practice while blocking agentification.


G. Evaluation Awareness Language

Use:

  • evaluation sensitivity
    (behavioral variance conditional on inferred oversight context)

Avoid:

  • “the model knows it is being evaluated”
  • “the model tries to pass the test”

H. Governance Boundary Rule

Language hygiene must preserve this invariant:

Improved observability ≠ governance
Behavioral compliance ≠ alignment
Diagnostic signals ≠ authority

Any analysis that blurs these boundaries requires revision.


I. Failure Mode Classification (New)

Ontological Slippage — ACP Failure Mode

Occurs when:

  • instrumental descriptors are interpreted as agent properties
  • models are treated as holders of intent or strategy
  • diagnostics are mistaken for control mechanisms

This failure mode must be flagged explicitly when detected.


J. Enforcement Guidance

  • This appendix governs:
    • ACP analyses
    • Ghost posts
    • policy-facing summaries
  • It does not prohibit exploratory internal modeling, but such modeling must be explicitly labeled non-canonical.