Pre-Phase-7 Risk Map for a Governed Reasoning System

1. Purpose

This note identifies the major failure modes ACP may encounter as it transitions from:

  • Phase 6: ingestion substrate, governance surfaces, Docking Harness
    to
  • Phase 7: live reasoning, module experimentation, claim generation, frontend exposure

The objective is not merely to predict likely software bugs.

It is to identify where ACP could fail as a knowledge institution.

That means this note considers:

  • engineering failure
  • ontology failure
  • governance failure
  • reasoning failure
  • interface failure
  • organizational failure

ACP is unusual because its failure is not measured only by crashes or bad outputs.

ACP can also fail by becoming:

  • untraceable
  • ungoverned
  • narratively persuasive but structurally weak
  • operationally convenient but constitutionally incoherent

2. Core Distinction

A conventional AI system fails when:

  • it hallucinates
  • it breaks
  • it becomes inaccurate

ACP can fail in those ways too.

But ACP also fails if it quietly becomes:

just another agentic text system

That is the deeper risk.

So the central question is:

Can ACP remain an artifact-governed knowledge substrate once reasoning turns on?

3. Primary Failure Classes

ACP’s risks fall into seven major classes:

  1. Evidence failures
  2. Reasoning failures
  3. Primitive / ontology failures
  4. Governance failures
  5. Docking Harness failures
  6. Module and UI failures
  7. Organizational / developmental failures

4. Evidence Failures

These are the most fundamental risks, because ACP begins with artifacts.

4.1 Artifact identity drift

The same source material produces different artifact identities across:

  • environments
  • ingestion paths
  • time
  • normalization states

This breaks:

  • deduplication
  • traceability
  • claim stability
  • long-term reproducibility

Why it matters

If artifact identity is unstable, the entire reasoning chain becomes unstable.

Design response

  • deterministic normalization
  • fixed hashing rules
  • golden artifact fixtures
  • cross-environment tests

4.2 Artifact mutation after ingestion

An artifact’s content changes after it is ingested.

This could happen through:

  • accidental edits
  • migration transforms
  • convenience updates
  • metadata confusion

Why it matters

ACP depends on the idea that artifacts are evidence surfaces.

If they mutate, then claims no longer point to stable evidence.

Design response

  • append-only artifact storage
  • no update path for content
  • explicit artifact-versioning if change is ever needed

4.3 Reasoning on unstable artifacts

Claims get generated from artifacts still in:

  • raw
  • triage

rather than stable.

Why it matters

ACP’s governance discipline depends on the difference between “present” and “ready for reasoning.”

Design response

Hard invariant:

no Claim / Scenario creation unless artifact_state >= stable

4.4 Provenance incompleteness

Artifacts exist, but their associated fields are weak or missing:

  • manifest
  • provenance
  • extracted text
  • source metadata
  • lineage context

Why it matters

An artifact without provenance is only partly ingested.

Design response

  • minimum artifact completeness contract
  • reject incomplete ingestion
  • provenance integrity checks

5. Reasoning Failures

These begin once Phase 7 starts generating claims and scenarios.

5.1 Claim without evidence

A claim is generated without traceable artifact support.

This is one of the clearest constitutional failures.

Why it matters

ACP’s basic sequence is:

Artifact → Claim → Scenario → DecisionEvent

A claim without an artifact breaks the chain at the start.

Design response

  • Claim creation must require artifact linkage
  • no free-floating claims
  • evidence boundary enforced at write time

5.2 Wrong-artifact grounding

A claim points to a real artifact, but the wrong one.

This is more subtle than no evidence, and probably more common.

Why it matters

The system appears grounded while actually being wrong.

This is a dangerous “looks legitimate” failure.

Design response

  • artifact-version binding
  • evidence spans / excerpt references where possible
  • review tools that expose claim-to-artifact mismatch

5.3 Scenario inflation

Scenarios begin to function as vague containers for ideas, notes, or synthesis, instead of structured reasoning objects.

Why it matters

This causes ontology drift and makes ACP less legible over time.

Design response

Scenarios must require:

  • bounded evidence set
  • explicit claim set
  • clear purpose

5.4 DecisionEvent laundering

DecisionEvents get created after the fact to make outputs look governed.

Why it matters

This turns governance into a cosmetic layer.

Design response

DecisionEvents must require:

  • actor identity
  • role
  • target object
  • timestamp
  • event legitimacy checks

5.5 Compression-induced false coherence

Reasoning outputs summarize disagreement too aggressively.

Divergent claims get flattened into a neat narrative.

Why it matters

This creates false consensus.

It is especially dangerous in governance or civic contexts because it makes the system appear more certain, aligned, or authoritative than it is.

Design response

  • preserve disagreement structurally
  • no default single-story synthesis
  • summaries must link back to underlying claim set

This risk is strongly echoed by adjacent research where longer or broader context can actually dilute signal and degrade answer quality rather than improve it.


6. Primitive / Ontology Failures

These are among the most dangerous long-term risks because they often look like productivity improvements.

6.1 Primitive creep

People introduce quasi-primitives such as:

  • Insight
  • Theme
  • Finding
  • Resolution
  • Recommendation
  • Cluster
  • Position

Why it matters

ACP’s discipline depends on a small grammar:

ArtifactVersion
Claim
Scenario
DecisionEvent

Once that grammar expands casually, the system loses coherence.

Design response

  • no new primitives without explicit governance
  • represent new concepts as metadata, views, or relationships over existing primitives

6.2 Frontend-driven ontology distortion

The frontend needs convenience models and begins to define the system more than the backend ontology does.

Why it matters

The UI can silently become the de facto ontology.

Design response

  • FE derives from primitives
  • no parallel frontend object model that becomes normative
  • UI labels must map back to canonical primitives

6.3 Registry creep

The module registry begins storing workflow logic, permissions, or hidden semantic categories instead of just module metadata.

Why it matters

This turns a minimal substrate into a logic dump.

Design response

  • registry remains metadata-only
  • all interpretive behavior lives outside the registry

7. Governance Failures

These are the failures that would most directly negate ACP’s purpose.

7.1 Governance becoming advisory again

Rules exist in docs or convention but are not enforced mechanically.

Why it matters

ACP has already moved through earlier phases precisely to avoid this.

Design response

Every critical invariant should live in:

  • code
  • tests
  • gate checks
  • runtime validation

not only prose


7.2 Promotion without governance objects

Artifacts or derivative reasoning outputs move up the lifecycle without explicit authorization records.

Why it matters

This bypasses institutional control.

Design response

Promotion requires:

  • governance object
  • target state
  • append-only log
  • refusal pathway

7.3 Refusal bypass

A refused artifact, claim, or scenario is reused downstream anyway.

Why it matters

Refusal becomes symbolic rather than operative.

Design response

  • refusal propagation as a hard invariant
  • downstream creation blocked if upstream object is refused

7.4 Exploratory-to-canonical shortcut

Experimental outputs become stable or canonical without re-docking / formal review.

Why it matters

This is one of the most likely “convenience corruption” patterns.

Design response

  • exploratory origin must be preserved
  • promotion requires explicit review path
  • no shortcut route into canonical status

7.5 Governance fatigue

Too many checks produce reviewer exhaustion, and humans begin rubber-stamping.

Why it matters

Governance can fail through volume, not only through absence.

Design response

  • priority banding
  • compact review queues
  • “no recommendation” should be valid
  • audit surfacing must be high signal

8. Docking Harness Failures

The Docking Harness is powerful, but it introduces a new control layer and therefore new failure risks.

8.1 Advisory-to-authority drift

Aalam or DH recommendations begin to function as de facto decisions.

Examples:

  • “safe to merge” starts being treated as merge authority
  • “close issue” starts being treated as closure
  • recommendations are not distinguished from decisions

Why it matters

This is perhaps the single biggest constitutional risk once DH is active.

Design response

  • all outputs explicitly labeled as recommendation / audit / proposal
  • no write path into merge / close / ratify
  • humans remain named authority

8.2 Role confusion between Atlas and Aalam

Atlas verifies.
Aalam reasons.
But over time they blur.

Why it matters

Separation of functions erodes.

Design response

  • actor-type contracts
  • distinct output types
  • DH enforcement of role boundaries

8.3 Hidden bypass paths

Governed DH exists, but old endpoints or direct service paths allow actions outside it.

Why it matters

The system has a constitution, but also side doors.

Design response

  • route inventory
  • CI checks for bypassable paths
  • deprecate legacy routes

8.4 GitHub-state contamination

Aalam scans old issues, stale PRs, superseded architecture notes, and treats them as current truth.

Why it matters

Repo history becomes accidental evidence.

Design response

  • freshness classification
  • distinguish active / stale / superseded / historical
  • recommendations must cite current state explicitly

8.5 Recommendation overload

DH produces so many closure / merge / refactor suggestions that humans stop reading carefully.

Why it matters

Audit becomes noise.

Design response

  • recommendation thresholds
  • triage ranking
  • bounded review batches

9. Module and UI Failures

These emerge quickly once Phase 7 begins.

9.1 Module drift

Modules begin acting like independent mini-systems rather than domain lenses over shared primitives.

Why it matters

ACP fragments.

Design response

Modules may vary in interpretation, but not in primitive structure.


9.2 Domain overfitting

One module, such as language or democracy, begins to dominate the system’s implied ontology.

Why it matters

ACP should remain a substrate, not a disguised single-domain product.

Design response

  • preserve primitive neutrality
  • domain logic stays in modules, not substrate

9.3 UI authority illusion

Frontend language implies that ACP has decided, resolved, ranked, or determined something.

Why it matters

Users will infer authority from UI wording even if the backend remains advisory.

Design response

Avoid labels like:

  • resolved
  • best answer
  • recommended outcome
  • system decision

unless backed by a real DecisionEvent


9.4 Agreement illusion

Visual aggregation makes plural claims look like one settled interpretation.

Why it matters

Plurality collapses into false consensus.

Design response

  • show disagreement
  • maintain provenance visibility
  • distinguish claim count from institutional agreement

10. Organizational / Developmental Failures

These matter especially because ACP is being developed by a very small team.

10.1 Architectural exhaustion

The system grows conceptually faster than the team can stabilize it.

Why it matters

A small team can invent fast, but maintenance debt accumulates invisibly.

Design response

  • prefer invariant hardening over feature expansion
  • phase discipline
  • architecture docs kept current

10.2 Tool-confidence illusion

AI assistance makes it easy to generate code, docs, and proposals faster than they can be deeply validated.

Why it matters

Aalam can accelerate development, but also accelerate fragility.

Design response

  • substrate-first discipline
  • explicit verification gates
  • slow promotion of architectural changes

10.3 Local coherence / global incoherence

Each piece makes sense in isolation, but the total system drifts.

Why it matters

This is a classic small-team risk in ambitious systems.

Design response

  • periodic architecture review
  • canonical docs tied to implementation
  • DH can eventually help inspect drift between code and doctrine

10.4 Loss of conceptual sharpness

ACP could gradually become framed as:

  • an agent platform
  • a knowledge graph product
  • a civic AI app
  • a research assistant

instead of governed knowledge infrastructure

Why it matters

A project can survive technically while losing its core identity.

Design response

  • keep canonical problem statement visible
  • phase notes should restate purpose
  • architecture review should include “what system is this becoming?”

11. Research Risks

These are not bugs in ACP, but risks in how ACP may relate to the external field.

11.1 Convergence risk

Other teams may independently arrive at:

  • claim-level auditability
  • provenance graphs
  • governed workflows

Why it matters

ACP’s distinctiveness may narrow over time.

Response

ACP should not rely on novelty alone; it should rely on coherence, execution, and institutional usefulness.


11.2 Legibility gap

ACP may be difficult for others to understand because it spans too many categories:

  • provenance
  • governance
  • RAG
  • knowledge graphs
  • agents
  • workflow systems

Why it matters

The project may be more original than it is communicable.

Response

  • architecture maps
  • phase docs
  • concise external descriptions
  • eventual paper or essay clarifying category

11.3 Premature category capture

If ACP is described too quickly using existing categories, outsiders may misunderstand it as:

  • RAG
  • multi-agent orchestration
  • deep research
  • knowledge graph QA

Why it matters

That would flatten its architectural ambition.

Response

When describing ACP, emphasize:

  • artifact-first
  • governed primitives
  • institutional memory
  • advisory-only AI roles

12. The Three Most Likely Early Phase-7 Failures

If I had to predict the first serious failures once Phase 7 starts, they would be these:

12.1 Grounding drift

Claims will be technically linked to artifacts, but often to the wrong artifact, wrong version, or unstable artifact state.

12.2 Advisory drift

Aalam / DH outputs will start to feel authoritative because they are useful and often correct.

12.3 Primitive drift

Modules and FE work will pressure the system to invent convenience objects outside the four primitives.

These are the ones to design against immediately.


13. Immediate Hardening Priorities

Before or at the start of Phase 7, the most valuable protections are:

  1. No claim/scenario unless artifact exists
  2. No claim/scenario unless artifact_state >= stable
  3. No external URL access in runtime reasoning
  4. No artifact mutation after ingestion
  5. No promotion without governance object
  6. No refusal bypass
  7. No exploratory-to-canonical shortcut
  8. No AI self-authorization
  9. No new primitives through modules or FE
  10. No consensus-looking compression of disagreement

14. Final Assessment

ACP’s greatest risk is not that it will fail like a normal AI application.

Its greatest risk is that it will partially succeed while slowly becoming something less governed, less traceable, and less constitutionally distinct than intended.

That means the central discipline for Phase 7 is:

Do not let reasoning outrun substrate integrity.

If ACP preserves that discipline, then Phase 7 can move fast without dissolving the architecture built in Phases 5 and 6.