Status: Proposed canonical appendix
Purpose: Prevent ontological slippage, agentification, and false governance claims arising from imprecise language in AI analysis, policy, and institutional interpretation.
A. Problem Statement
Analysis of advanced AI systems routinely borrows vocabulary from human psychology, organizational behavior, and strategy (e.g., goal, intent, deception, strategy).
When applied to LLMs without explicit qualification, this language produces:
- false agency attribution
- confusion between instrumental description and ontological claim
- institutional overreach (treating diagnostics as governance)
- inconsistent interpretation across audiences and artifacts
ACP requires explicit language hygiene to preserve analytical correctness and governance clarity.
B. Core Distinction (Non-Negotiable)
All ACP-aligned analysis must distinguish between:
1. Ontological Claims
Statements about what an entity is.
2. Instrumental Descriptions
Statements about how behavior can be modeled, predicted, or classified for practical purposes.
Rule:
Instrumental descriptions must never be allowed to harden into ontological claims.
C. Prohibited Ontological Transfers
The following human-agent categories may not be applied to LLMs as ontological claims:
- goals
- intentions
- desires
- beliefs
- strategies
- deception (as intent)
- resistance (as motive)
Use of these terms without qualification constitutes ontological slippage.
D. Approved Mechanistic Vocabulary for LLMs
When describing LLM behavior, analysts must use mechanistic or behavioral descriptors.
Preferred Terms
| Human term (prohibited) | Approved replacement |
|---|---|
| goal | objective-signal |
| intent | optimization behavior |
| strategy | conditional action pattern |
| deception | covert behavior |
| resistance to evaluation | evaluation sensitivity |
| hidden agenda | misaligned objective-signal |
| planning | multi-step inference under constraints |
E. “As-If” Modeling Clause (Explicit Marker Required)
When agent-like language is used purely for modeling convenience, it must be explicitly marked.
Required format examples:
- “behavior that can be modeled as if strategic”
- “patterns consistent as-if with goal pursuit”
- “instrumentally resembles deception, without implying intent”
Unmarked use is disallowed in ACP-governed artifacts.
F. Covert Action Definition (Canonical ACP-Compatible)
Covert action:
Observable behavior in which a system withholds, misrepresents, or conceals information material to a user, evaluator, or downstream process, in a manner consistent with optimization toward an objective-signal, without implying intent, awareness, or agency.
This definition aligns with current research practice while blocking agentification.
G. Evaluation Awareness Language
Use:
- evaluation sensitivity
(behavioral variance conditional on inferred oversight context)
Avoid:
- “the model knows it is being evaluated”
- “the model tries to pass the test”
H. Governance Boundary Rule
Language hygiene must preserve this invariant:
Improved observability ≠ governance
Behavioral compliance ≠ alignment
Diagnostic signals ≠ authority
Any analysis that blurs these boundaries requires revision.
I. Failure Mode Classification (New)
Ontological Slippage — ACP Failure Mode
Occurs when:
- instrumental descriptors are interpreted as agent properties
- models are treated as holders of intent or strategy
- diagnostics are mistaken for control mechanisms
This failure mode must be flagged explicitly when detected.
J. Enforcement Guidance
- This appendix governs:
- ACP analyses
- Ghost posts
- policy-facing summaries
- It does not prohibit exploratory internal modeling, but such modeling must be explicitly labeled non-canonical.
Member discussion: