Purpose
This document supplements Substrate Constraint Canon v4. It does not replace it. Canon v4 defines the core laws, ontology, promotion model, evaluation engine, validation modes, and domain inheritance structure for Substrate. This supplemental document preserves the operational material that emerged around those laws: failure pressure, proof discipline, enforcement reality, PR governance, system state grounding, and the transition logic from MVP loop to future architecture.
The reason this document exists is practical. The working materials that produced Canon v4 contain important execution knowledge that should not be lost when earlier documents are archived. Some of that material is redundant with Canon v4 and should not be repeated in full. Some is obsolete because PR7 works and PR8/PR9 have already merged. Some is not yet active but must be preserved for later stages. The function of this supplement is to consolidate what remains valuable: the material that explains how the canon behaves in practice, where the system is most likely to fail, what must be proven before expansion, and which operational rules should shape the next stage of development.
The intended reader is the system owner. This document is not written primarily as an engineer ticket, external explainer, or public manifesto. It is a steering document. It should help determine what is already stable, what must be watched, what must not be prematurely built, and what should eventually be folded into a future Canon v5.
Canon v4 answers: What is a constraint, and how does it govern?
This supplement answers: What breaks, where enforcement actually lives, and what proof is required before the system may advance?
Reader Roles
System owner: use this to decide whether to proceed, stop, narrow, or defer.
Aalam instance: use this to enforce proof discipline, reject imaginary system state, and preserve failure.
Engineer: use this to scope PRs, prove behavior, avoid future leakage, and separate FE/BE confidence.
Future domain lead: use this to inherit constraints without prematurely building domain systems.
Use Parts I–XI as the primary steering narrative.
Use Addenda A–K as registries and operational references.
For build review, start with Part V, Part VI, Addendum E, and Addendum F.
For system design, start with Part VIII, Addendum C, Addendum D, and Addendum K.
For future Canon v5 synthesis, use Part X, Addendum A, and Addendum J.
Status: Final Supplemental Draft for Stage 0 / 7.1. Authoritative for project steering and onboarding; not a replacement for Canon v4; candidate source for Canon v5.
Part I — Relationship to Canon v4
Canon v4 is the primary canonical reference. Its core structure should remain stable unless future evidence requires revision. It defines constraints as system-defining conditions that determine what may exist, proceed, be promoted, be blocked, or be preserved as failure. It also establishes the core lifecycle: signal → claim → artifact → evidence → validation → promotion → enforcement → monitoring → revision.
This supplement does not restate that canon in full. Instead, it extends Canon v4 in four ways.
First, it preserves execution rules that were developed during the MVP loop period. These include slice isolation, proof over plausibility, negative validation, no scope contamination, and the requirement that each PR represent one minimal user-visible capability. These rules began as MVP-specific but remain valuable as project discipline.
Second, it consolidates failure material. The uploaded failure taxonomy identifies epistemic failures, reasoning failures, execution failures, system design failures, organizational failures, human-level failures, market failures, and governance failures. These do not all belong in Canon v4 as independent system laws, but they are important because they explain why the canon exists and where Substrate is likely to fail under pressure.
Third, it preserves proof and validation standards. The evaluation protocol and proof register distinguish technical completion from system proof. A working loop is not enough; reuse must be explicit, visible, comparative, and materially better than baseline chat.
Fourth, it marks future material as deferred rather than discarded. Full claim schema, trace system, decision-event schema, governed artifact lifecycle, federation, ingestion pipelines, advanced UI, and detailed domain builds are not appropriate for the supplemental body. They belong in the future-work document that follows this one. Their omission from this supplement is not rejection.
Part II — Supplemental Constraint Extensions
The following constraints are not replacements for Canon v4. They are extensions generated by execution, validation, and failure analysis. They should remain supplemental for now, but they are candidates for synthesis into a future Canon v5 after additional testing.
Part II introduces the highest-priority supplemental constraint extensions in narrative form. Addendum A preserves the complete constraint-extension registry, including additional candidates, status notes, source clusters, and future Canon v5 implications.
SC-01 — Structure Must Do Work
A structure is valid only if it changes system behavior, enables validation, constrains output, improves reuse, or reduces failure. A schema, label, field, checklist, or artifact type that merely organizes text is not yet system structure. It may be useful as documentation, but it does not govern.
This matters because Substrate is especially vulnerable to fake rigor. The system can produce documents, schemas, issue maps, stage maps, and validation tables that look disciplined while adding no control. The relevant question is not whether a structure is clear. The relevant question is what the structure prevents, enables, tests, or changes.
Failure signal: structured output appears more rigorous than ordinary chat, but downstream behavior is unchanged.
SC-02 — Proof Over Plausibility
A PR, artifact, validation run, or system claim is not accepted because it appears plausible. It requires explicit proof. In the MVP execution docs, valid proof included DB evidence, test output, API response, or observable system behavior. UI appearance, code presence, expectation, or developer confidence were explicitly insufficient.
This rule remains important beyond the MVP. Substrate repeatedly risks confusing visible structure with demonstrated mechanism. Proof over plausibility blocks that drift.
Failure signal: a capability is treated as existing because the UI suggests it, because code exists, or because the architecture intends it.
SC-03 — Negative Validation Is Required
A system slice is not proven by showing only that the intended path works. It must also show that excluded behavior does not occur. In the PR4–7 execution docs, if reuse was out of scope, reuse had to be inactive; if retrieval was out of scope, retrieval could not be partially active; if future hooks existed, they had to be inert.
This rule should remain active because many system failures appear through leakage. A feature can appear correctly scoped while carrying hidden future behavior. Negative validation protects sequencing.
Failure signal: a PR passes positive tests but contains hidden future logic, implicit reuse, partial automation, or untested excluded behavior.
SC-04 — Reuse Must Be Causal
Reuse is real only if the reused artifact materially changes the output in a way that can be attributed to the artifact. The MVP validation addendum states that good reuse requires contextual fit, structural leverage, transformation capability, outcome improvement, reduced effort, and consistency over time.
This rule is central. Artifact reuse is the core mechanism that justifies Substrate. If artifacts are stored and retrieved but do not improve reasoning, the system is only a managed note store. If reuse appears to improve output but the improvement is caused by hidden context, model capability, or restated user instructions, the system has not proven its primitive.
Failure signal: the output looks better, but the artifact’s contribution cannot be isolated.
SC-05 — Visibility Defines Operational Reality
A capability that is not visible, user-triggerable, and inspectable should not be treated as active. The current-system-state document states the principle directly: what is not implemented and user-visible does not exist.
This does not mean every backend operation must be shown in detail. It means that system-relevant state must be visible at the point where it affects reasoning or authority. Artifact creation, retrieval, reuse, validation status, failure state, and decision boundaries cannot be hidden without undermining governance.
Failure signal: the system claims reuse, validation, memory, or authority, but the user cannot see where it enters the workflow.
SC-06 — Slice Isolation Protects Causal Learning
Each PR or build unit must represent one minimal, independent capability. It must define what it does, what it does not do, where the proof is, and what happens when the system is tested against its boundaries. The MVP governance docs treated this as non-negotiable because the point was not merely to ship features but to learn whether the mechanism works.
Slice isolation is especially important because Substrate’s phases build on each other. If a slice contains hidden dependencies or future behavior, later results cannot be interpreted. The system may work, but the reason it works will be unclear.
Failure signal: a later capability appears to succeed because an earlier PR quietly included part of it.
SC-07 — Technical Completion Is Not System Proof
A technically functioning loop does not prove the system hypothesis. The evaluation protocol explicitly separates implementation readiness from evaluation success: evaluation begins only after the full loop works, but the loop passes only if reuse produces visible, repeatable reasoning value over baseline chat.
This distinction prevents demo theater. A working UI, persistent artifacts, retrieval, and reuse are prerequisites. They are not the result. The result is comparative reasoning improvement.
Failure signal: the project advances because the loop executes, even though reuse has not demonstrated material advantage.
SC-08 — Caveated Validation Is Valid Only When the Caveat Is Bound
Some results are neither clean pass nor failure. The caveated validation rule allows progress when the core mechanism is sufficiently proven and remaining weaknesses are explicit, bounded, and assigned to follow-up. The caveat must not undermine the core conclusion, and it must not be presented as stronger evidence than it is.
This rule matters because strict binary evaluation can create false blockers, while loose caveating can create false confidence. The middle path is honest bounded validity.
Failure signal: a limitation is known but minimized, or a partial result is treated as full validation.
SC-09 — Signal Must Be Preserved Under Expansion
Expansion is valid only if causal signal remains attributable. If adding features, users, domains, agents, ingestion, or governance layers makes it unclear why the system works or fails, expansion has outrun proof. This rule emerged from validation runs in which structured reuse and variation produced value only when the mechanism remained inspectable.
This constraint protects against premature scale. It does not prohibit expansion. It requires that expansion preserve the ability to observe causal effect.
Failure signal: system complexity increases and improvement can no longer be attributed to artifact reuse, validation, variation, or another named mechanism.
SC-10 — Insight Requires Structured Variation, Not More Output
The validation materials identify an important extension of reuse: insight can emerge from structured variation, comparison, and synthesis. In one validation sequence, reuse plus variation produced new conceptual structure that did not exist in any single output.
This is important for future Substrate vision. The system is not only a stabilizer or memory layer. Under the right conditions, it can become an insight-generation system. But variation is not a default mode. It should be triggered when a single reasoning path is insufficient, uncertainty is high, competing explanations exist, or the user needs stronger confidence.
Failure signal: the system generates multiple variants that are redundant, superficial, or ungoverned.
Part III — Failure as First-Class System Input
Failure is not merely an undesirable result. In Substrate, failure is the main source of new primitives, constraints, validation methods, and promotion criteria. The failure taxonomy shows that AI systems do not fail only because they hallucinate. They fail because outputs appear correct before they are verified, actions become possible before they are constrained, and knowledge spreads faster than it can be validated.
Canon v4 already states that failure must be preserved. This supplement adds the operational interpretation: failure must be mapped to the missing primitive or weak enforcement layer that allowed it.
Epistemic Failure
Epistemic failures occur when the system produces or preserves false, unsupported, misleading, or overconfident knowledge. The taxonomy includes hallucination, mirage effects, phantom competence, epistemic contamination loops, and synthetic consensus.
The required primitive response is claim isolation plus validation. Compound outputs must be decomposed into units that can be supported, contradicted, deferred, or rejected. Agreement among models or agents cannot substitute for evidence.
Operational implication: PR9+ artifact review must ask whether an artifact contains claims, whether those claims are isolated, and whether evidence or validation mode is declared.
Reasoning Failure
Reasoning failures occur when outputs imitate reasoning without reliable underlying logic. The taxonomy identifies pattern imitation, distilled error propagation, overgeneralization, and anthropomorphic projection.
The required primitive response is structured decomposition. A reasoning artifact must preserve assumptions, steps, constraints, uncertainty, and decision logic where relevant. The goal is not to force every output into a rigid claim schema, but to prevent blob reasoning from being treated as validated reasoning.
Operational implication: artifacts should preserve reusable structure, not merely polished conclusions.
Execution Failure
Execution failures occur when systems act, mutate state, or imply authority without sufficient validation. The taxonomy identifies fail-open execution, ambient delegation, structured execution without correctness, authority escalation, and hidden state mutation.
The required primitive response is decision-event and enforcement architecture. These are not active now in full form, but their necessity is already visible. The system must eventually distinguish proposal from approval, approval from execution, and execution from verification.
Operational implication: current systems should avoid execution authority, and future systems must not connect actions to outputs without decision gates.
System Design Failure
System design failures occur when tools are mistaken for systems. The taxonomy identifies tool ≠ system, demo illusion, strategy without process audit, and discovery without constraint.
The required primitive response is the docking harness: a forced pathway that prevents outputs from bypassing artifacts, validation, trace, and enforcement. This remains future work, but the principle applies now: every output that matters must move through the governed loop.
Operational implication: a polished demo is not evidence unless it proves the mechanism under real or comparable conditions.
Organizational and Human Failure
Substrate must account for the fact that human oversight is not automatically reliable. The governance-problems doc identifies weak HITL, cognitive overload, time pressure, unclear responsibility, and superficial approval as core failure patterns. The failure taxonomy adds cognitive debt, orchestration without understanding, overtrust, and talent-pipeline degradation.
The required primitive response is not “more human review.” It is structured human authority: explicit checkpoints, visible evidence, bounded responsibility, and artifacts that allow review to be meaningful.
Operational implication: human involvement must be designed, not assumed.
Meta-Failure
The central meta-failure is capability-control mismatch. Systems can generate, act, scale, and persuade before they can verify, constrain, or understand their own outputs. The failure taxonomy compresses this as: the system produces outputs, actions, and knowledge faster than it can verify any of them.
Substrate’s purpose is to slow that propagation by forcing outputs into artifacts, artifacts into validation, validation into constraints, and constraints into enforcement.
Part IV — Failure / Pressure Register
The Failure / Pressure Register should remain part of the supplemental because it is not merely historical. It identifies the pressures most likely to invalidate the system as it grows. The register should be updated only when a failure is newly observed, materially reinterpreted, mitigated, or shown to be weaker or stronger than expected.
The register requires one amendment from the working notes: each failure should include Stage Relevance and Blocks Progression. A fatal unresolved risk at a relevant stage blocks advancement.
FP-01 — Weak Artifact Degradation
Artifacts may degrade into weakly structured text, causing retrieval and reuse to operate over plausible but low-value objects. This is fatal because the entire loop depends on artifacts being more than stored chat fragments. If artifacts do not preserve reusable reasoning, the system stores noise.
Stage relevance: PR9.5, PR10, PR11, Stage 1.
Blocks progression: yes, if artifact quality remains unresolved.
FP-02 — Minimal Compliance / User Bypass
Users may comply with structure only superficially under time pressure, fatigue, or convenience incentives. This is fatal because Substrate can become orderly-looking process without epistemic gain.
Stage relevance: all stages, especially early external testing.
Blocks progression: conditionally, if real users create artifacts that satisfy form but not function.
FP-03 — Reuse Does Not Improve Outcomes
Artifact reuse may fail to improve clarity, continuity, correctness, usefulness, or effort over baseline chat. This is the single most important proof condition. If reuse does not materially improve outcomes, Substrate collapses into structured inefficiency.
Stage relevance: PR10, PR11, Stage 0 exit.
Blocks progression: yes.
FP-04 — Provenance Ignored
Provenance may exist but fail to affect interpretation, trust, or decision-making. This creates symbolic rigor without actual governance.
Stage relevance: trace, governance, frontend visibility, future shared systems.
Blocks progression: later, when provenance becomes active.
FP-05 — Wrong-but-Plausible Retrieval
Retrieval may return a plausible but wrong artifact, and users may not detect the mismatch. This is dangerous because the system fails cleanly and convincingly.
Stage relevance: retrieval, reuse, Stage 2+.
Blocks progression: yes, once retrieval quality is central.
FP-06 — Ambiguous Artifact State
Artifact state labels may become inconsistent or disconnected from actual behavior. If status labels do not affect reuse conditions or trust, governance becomes cosmetic.
Stage relevance: Stage 1+, governance stages.
Blocks progression: yes for any status-governed workflow.
FP-07 — Corpus Degradation
Artifact accumulation may create stale, contradictory, duplicated, or weakly maintained corpora. Memory without lifecycle discipline becomes clutter and increases false confidence.
Stage relevance: volume tests, ingestion, shared spaces.
Blocks progression: eventually, especially before ingestion or external use.
FP-08 — Cognitive Overhead Exceeds Value
Substrate may be too cumbersome relative to the benefit. Even if architecture is sound, overhead can kill adoption or restrict the system to narrow high-discipline users.
Stage relevance: all stages.
Blocks progression: if users do not voluntarily reuse artifacts or continue workflows.
FP-09 — Local Correctness Does Not Produce System Correctness
Individual artifacts may be locally valid while the system produces globally misaligned outcomes. This becomes more serious as artifacts interact, accumulate, and drive decisions.
Stage relevance: claim, trace, governance layers.
Blocks progression: later, especially before action or institutional deployment.
FP-10 — Governance Theater
Visibility, approval, acknowledgment, and status can become rituals that signal rigor without improving judgment.
Stage relevance: frontend visibility, authority, governance UI.
Blocks progression: once governance surfaces are introduced.
FP-11 — Premature Schema Path Dependency
Early schema or artifact design choices may constrain later governance awkwardly. This is moderate rather than fatal because minimalism can mitigate it, but it should be watched before activating later layers.
Stage relevance: schema design, claim storage, trace systems.
Blocks progression: no, unless schema decisions begin constraining validation or enforcement.
FP-12 — Narrow User Fit
The system may work only for reflective users, high-discipline workflows, or high-stakes reasoning. This may be a valid boundary rather than a defect, but it must be understood honestly.
Stage relevance: product positioning, language, institutional use.
Blocks progression: no, unless product claims exceed actual fit.
FP-13 — Structured Clarity Mistaken for Epistemic Improvement
Cleaner outputs may be mistaken for better outcomes. This is one of the strongest false-positive risks because Substrate’s structure can make weak reasoning look more trustworthy.
Stage relevance: PR9+, validation, external demo.
Blocks progression: yes, if baseline comparison is weak.
FP-14 — Premature External or Organizational Use
External or organizational use may begin before the core loop is empirically proven. This can generate noisy feedback, misframe the product, and create false judgments about value.
Stage relevance: external testing, accounts, institutions.
Blocks progression: yes, until Stage 0 exit is valid.
Part V — Execution Reality: How the Canon Survives Build
The purpose of this section is to preserve the execution doctrine that emerged from the MVP loop and PR governance documents. Canon v4 defines what constraints are and why they matter. This section defines how those constraints survive implementation. The distinction is important because a correct canon can still fail if the build process allows hidden state, forward leakage, weak proof, or ambiguous completion.
Substrate is not built by adding features. It is built by proving one constrained capability at a time. Each PR is therefore not merely a code change; it is a system-boundary event. It changes what the system can do, what may be tested, what may be trusted, and what later work is allowed to assume. If a PR is unclear, contaminated, or unproven, the damage is not local. It corrupts the interpretation of later stages.
The operating rule is that every build slice must be isolated, visible, provable, and bounded. A slice must implement exactly one capability, expose that capability through observable behavior, prove that the capability works, and prove that excluded capabilities do not work or do not exist. This is how the system preserves causal learning. Without slice isolation, a later success may be caused by hidden logic introduced earlier. Without negative validation, future behavior may leak into the current stage. Without proof, the project mistakes intention for implementation.
This execution model came from the PR4–7 MVP discipline, but it remains valid after that period. The original MVP loop required chat, artifact creation, storage, retrieval, and reuse to be built sequentially because the project needed to learn whether reusable artifacts had value. That same logic applies to later stages. Claim systems, trace systems, decision events, governance layers, docking harnesses, and federation must each be introduced only when the current stage has produced the evidence required to justify them.
A valid PR must therefore answer four questions clearly: what does it do, what does it explicitly not do, where is the proof, and what happens when the boundary is tested. If any answer is unclear, the PR is not ready. This is not bureaucracy. It is the mechanism that keeps Substrate interpretable.
Core Execution Requirements
Each PR must satisfy the following requirements before it can be treated as valid.
| Requirement | Meaning | Failure if missing |
|---|---|---|
| Slice isolation | One minimal capability only. | Causal contamination. |
| Defined exclusion | Clear statement of what is not included. | Future leakage. |
| Proof over plausibility | Observable evidence, not expectation. | Demo illusion. |
| Negative validation | Out-of-scope behavior tested inactive. | Hidden coupling. |
| Reproducibility | Another actor can verify the behavior. | Non-transferable proof. |
| Clean boundary | No unrelated or future-stage behavior. | Uninterpretable result. |
Valid proof includes database evidence, API responses, persisted state, test output, visible system behavior, or reproducible interaction. Invalid proof includes UI appearance alone, code presence alone, architectural intention, developer confidence, or a plausible explanation of what should happen. A system can look correct while failing to persist state. A feature can exist in code while not being user-accessible. A UI can imply a backend state that does not exist. Therefore, proof must attach to behavior, not to appearance.
Negative validation is equally important. A PR must demonstrate not only that the included capability works, but also that excluded behavior is inactive. If reuse is not in scope, attempts to trigger reuse should fail or do nothing. If automation is not in scope, no hidden automation should occur. If claim extraction is not yet active, outputs should not be treated as claims. This protects staged development from contamination.
The most common execution failures are predictable. UI illusion occurs when an interface appears functional but backend persistence or state integrity is missing. Hidden coupling occurs when future logic exists but is described as unused or preparatory. Implicit behavior occurs when the system performs a relevant operation without explicit user action. Partial features masquerade as complete when the happy path works but boundaries fail. Commit contamination occurs when unrelated or future-stage work enters a PR and makes the result difficult to interpret.
The practical command is simple: reject unclear work early. A PR that cannot prove its slice should not merge. A PR that includes future behavior should be split, removed, or made inert. A PR that has a strong backend but weak frontend, or strong frontend but weak backend, should not be evaluated as a single confidence state. Backend and frontend proof must remain separate where their evidence differs.
This execution layer is the first place where Canon v4 becomes operational. The canon says constraints must have consequence. PR execution is where that consequence begins.
Part VI — Proof as Decision Engine: Evaluation Determines Whether the System May Advance
The purpose of proof is not to show that the system can be demonstrated. Proof determines whether Substrate may continue to expand. At the current stage, the central hypothesis is narrow: explicit creation, persistence, retrieval, and reuse of artifacts must materially improve reasoning outcomes over baseline chat. Nothing else justifies expansion.
The MVP loop can be technically complete and still fail. A user may be able to create an artifact, store it, retrieve it, and reuse it, while the result remains no better than ordinary chat. In that case, the implementation has succeeded but the system hypothesis has not. This distinction is essential. Technical completion proves that a workflow exists. System proof requires that the workflow produces value.
Evaluation begins only after the full loop exists. The system must allow a user to create an artifact from real chat output, persist it across a session boundary, retrieve it later, explicitly reuse it in a new interaction, and see that reuse at the point of interaction. If any of these conditions are missing, evaluation has not begun. The correct response is not to interpret partial results; it is to finish or repair the loop.
Once the loop exists, evaluation must compare two conditions: baseline chat and artifact reuse. The baseline condition uses no artifact and no reuse. The reuse condition explicitly inserts the artifact into the reasoning context. The task must be the same or materially equivalent. Without comparison, improvement cannot be attributed. Without attribution, reuse is not proven.
The evaluation dimensions are clarity, faithfulness to prior reasoning, effort reduction, continuity across the session boundary, and reuse visibility. These dimensions matter because Substrate does not claim merely to produce prettier outputs. It claims to reduce repetition, preserve reasoning, improve continuity, and make prior work active in future thinking. If those effects are not visible, the loop is not validated.
Evaluation Gate
The loop can pass only when all of the following are true.
| Gate | Requirement |
|---|---|
| Full loop | Chat → artifact → storage → retrieval → reuse works end-to-end. |
| Explicit reuse | The user selects or invokes the artifact. |
| Visible reuse | The artifact’s presence is visible in the interaction. |
| Comparative gain | Reuse improves at least two evaluation dimensions. |
| Attribution | The improvement is plausibly caused by the artifact. |
| Repeatability | The result is not a one-off lucky demo. |
Unclear results do not count as success. In a prove-or-stop phase, weak evidence is not enough to justify expansion. If the result is mixed, ambiguous, fragile, or dependent on interpretation, the operational treatment is failure or redesign. This does not mean the system is worthless. It means the current proof is insufficient.
The Proof / Validation Register exists to prevent retrospective optimism. It breaks the proof problem into observable components: end-to-end loop existence, reuse visibility, cross-session persistence, baseline versus reuse comparison, user preference, and artifact usefulness. Each component can be scored, recorded, and revisited. The register prevents future Aalam instances or project participants from relying on memory, tone, or enthusiasm.
Proof records should become artifacts. Each evaluation should record the task, artifact used, baseline summary, reuse summary, comparison dimensions, final judgment, and required follow-up. This matters because proof itself must be durable. If the system is to build persistent reasoning artifacts, then its own validation history must also be preserved.
The critical insight is that a system that works but does not improve reasoning has failed the current hypothesis. It may still be a useful prototype, but it has not earned expansion into more complex architecture. The correct response is not to add claims, trace, governance, domains, or ingestion. The correct response is to redesign artifact quality, reuse mechanics, or interaction flow until causal improvement is visible.
This is why proof functions as a decision engine. It determines whether the system proceeds, pauses, narrows, or stops.
Part VII — Operating Discipline: Prove-or-Stop as System Method
The operating doctrine is the discipline layer that prevents Substrate from becoming its own failure mode. The broader vision is large: governed reasoning, artifacts, claims, trace, validation, decision events, docking harnesses, federation, domain modules, language, legislative drafting, institutional memory. That vision creates constant pressure to expand. The doctrine exists to resist that pressure until the current mechanism is proven.
The core principle is that the goal is not to build Substrate. The goal is to determine whether Substrate should exist in the form currently imagined. This principle was created for the MVP phase, but it remains useful throughout staged development. Each stage must earn the next stage. The existence of a future architecture does not justify building it.
At any given time, there is one active objective. During the MVP and current Stage 0 boundary, that objective is to prove the loop: chat → artifact → storage → retrieval → reuse → improved output. Work that does not directly support the active objective is out of scope. This includes interesting, plausible, or eventually necessary work. In Substrate, timing matters. A correct feature built too early can still be wrong because it destroys causal clarity.
End-to-end behavior is the standard. A feature is not complete because a backend model exists, because a frontend component exists, or because a prompt format works. It is complete only when a user can perform the relevant flow. Partial states do not count. Artifact storage without retrieval is not the loop. Retrieval without reuse is not the loop. Reuse without observable effect is not the loop. Technical completion without system value is not the loop.
Manual behavior is preferred before automation because manual behavior preserves causality. In early stages, the user should explicitly select artifacts, explicitly reuse artifacts, and explicitly observe the result. Automatic retrieval, automatic injection, hidden memory, and implicit context may eventually be useful, but they are dangerous before the mechanism is proven. Automation hides the very causal paths the system needs to test.
No work should be justified by future architecture. The correct question is not whether a feature will eventually be needed. The correct question is whether it is required for the current stage’s proof. If the answer is no, it should be deferred. This is especially important for claim schemas, trace systems, decision events, governance workflows, ingestion pipelines, and multi-agent systems. All may become important. None should be allowed to contaminate current proof.
When uncertainty arises, the default is elimination rather than expansion. Remove scope, simplify the slice, delay the feature, or reduce the claim. This is not minimalism for its own sake. It is signal preservation. Complexity expands faster than validation. The system should grow only when evidence forces growth.
Operating Rules
| Rule | Operational meaning |
|---|---|
| Single objective | Only the active proof target matters. |
| End-to-end only | Partial capability does not count. |
| Manual before automated | Explicit action preserves causal visibility. |
| No hidden systems | If it cannot be seen or triggered, it does not exist operationally. |
| No future justification | Future value does not justify current build. |
| Eliminate before expanding | Uncertainty should reduce scope. |
| Strict sequencing | Later layers wait until earlier proof exists. |
| Stop on failed mechanism | Failure triggers redesign, not feature growth. |
A weekly or stage-level reality check should ask whether a real user can complete the current loop and whether that loop produces the expected value. If the answer is no, development remains on the current mechanism. If the answer is yes, evaluation begins. If evaluation fails, expansion stops.
This doctrine is not anti-ambition. It is what protects the ambition from collapsing under premature complexity.
Part VIII — System State Reality: What Exists, What Does Not, and Why It Matters
A central risk in Substrate is reasoning from imagined architecture rather than implemented reality. The project has generated extensive designs for claims, trace, decision events, constraints, governance, domain systems, institutional roles, and federation. Those designs are valuable, but they must not be treated as active system capability.
The current reality is narrower. The active or near-active system consists of chat, artifacts, storage, retrieval, and reuse. The minimal MVP data model requires user, chat session, chat message, and artifact. Every artifact must have an owner and must be retrievable only by that owner. The minimal loop is sufficient if one user can send a chat message, receive a response, create an artifact from that response, save it, list prior artifacts, reopen one later, and reuse it in a new interaction.
The current system does not yet include a full claim layer, full trace system, decision-event layer, enforcement engine, federation system, ingestion pipeline, multi-user institutional workspace, or governed domain module. These items are future systems. They may be conceptually designed, but they are not active unless implemented, visible, and tested.
This distinction matters because imaginary layers can contaminate evaluation. An output should not be interpreted as a validated claim if no claim layer exists. A stored artifact should not be treated as traceable provenance if no trace system exists. A user approval should not be treated as a decision event if no decision-event structure exists. A UI label should not be treated as governance if it does not alter behavior. A future architecture cannot be used as evidence for current system capability.
Current-State Boundary
| Category | Current status | Operational rule |
|---|---|---|
| Artifact | Active / core primitive | Can be created, stored, retrieved, reused. |
| Reuse | Active / under proof | Must be explicit, visible, causal. |
| Claim | Conceptual / partial lens | Do not treat outputs as structured claims yet. |
| Trace | Minimal source linkage only | Do not claim full provenance. |
| Decision event | Deferred | Do not treat outputs as authorized decisions. |
| Enforcement engine | Deferred | Process enforcement is not system enforcement. |
| Federation | Deferred | No distributed governance yet. |
| Ingestion | Deferred / future | Do not bulk-parse before artifact quality and signal controls. |
The reality rule is: what is not implemented and user-visible does not exist as system capability. This should not be softened. It protects the project from hallucinating its own infrastructure.
The practical consequence is that all current reasoning must be grounded in what the system can actually do. Future systems should be preserved in roadmap documents, not smuggled into current-stage interpretation.
Part IX — Failure Response: From Observation to System Change
Failure is not complete when it is noticed. In Substrate, a failure becomes useful only when it produces one of three outcomes: constraint refinement, validation strengthening, or system redesign. A failure that is merely described remains an observation. A failure that changes system behavior becomes governance input.
The failure register identifies several fatal or serious pressures: weak artifact degradation, minimal compliance, reuse failure, ignored provenance, wrong-but-plausible retrieval, ambiguous artifact state, corpus degradation, excessive cognitive overhead, local correctness without system correctness, governance theater, premature schema path dependency, narrow user fit, structured clarity mistaken for improvement, and premature external use. These are not all equally urgent, but each must eventually connect to an action path.
Artifact failure requires artifact design changes. If artifacts degrade into weak structured text, the response is not to improve retrieval. It is to strengthen artifact criteria, reject weak artifacts, preserve problem type and decision logic, and test reuse quality. Retrieval over weak artifacts only retrieves weakness more efficiently.
Reuse failure is a stop condition. If artifacts can be created and retrieved but do not improve outcomes, the system should not expand. The response is to redesign reuse mechanics, compare baseline and reuse, examine artifact quality, and retest. Adding domains, accounts, ingestion, or agents would only add complexity to an unproven mechanism.
User compliance failure requires interaction redesign. If users fill structures minimally, skip artifacts, or perform schema completion without meaning, the system must reduce cognitive load, improve artifact affordances, make value visible, or narrow the user profile. The response is not to blame the user; it is to examine whether the system asks for too much structure before delivering enough value.
Wrong-but-plausible retrieval requires retrieval validation. The system must eventually justify why an artifact is relevant, expose the artifact at the point of reuse, and allow rejection. This becomes increasingly important as artifact volume grows.
Governance theater requires binding status to behavior. A badge, approval label, provenance marker, validation state, or authority flag must either alter admissibility, reuse, promotion, or decision rights, or be clearly marked as informational. Otherwise the system trains users to trust decorative governance.
Structured clarity mistaken for improvement requires comparative evaluation. Cleaner output is not necessarily better reasoning. The system must compare against baseline and identify whether clarity, faithfulness, effort, or continuity actually improved.
Failure Response Table
| Failure | Required response | Stop condition |
|---|---|---|
| Weak artifacts | Strengthen artifact criteria; reject weak artifacts. | If artifacts cannot support reuse. |
| Reuse failure | Redesign reuse; rerun comparison. | If reuse does not improve outcomes. |
| Minimal compliance | Redesign interaction; reduce overhead. | If users do not naturally reuse artifacts. |
| Wrong retrieval | Add relevance checks and visibility. | If wrong artifacts are reused undetected. |
| Governance theater | Bind status to behavior. | If labels imply false authority. |
| Structured clarity illusion | Force baseline comparison. | If improvement is only aesthetic. |
| Scope contamination | Split or reject PR. | If causal interpretation is lost. |
| Future leakage | Remove or make inert. | If current proof depends on future logic. |
Failure should not automatically create more architecture. Sometimes the correct response is to remove scope. Sometimes it is to strengthen validation. Sometimes it is to stop. This is why failure preservation matters: not because the system collects mistakes, but because preserved failure determines what the system must become next.
Part X — Constraint Promotion as Authority Transfer
Promotion is the mechanism by which observations, rules, failures, and constraints gain authority. It is not a reward for usefulness. It is not a label for ideas that seem important. Promotion transfers power inside the system. A promoted constraint can affect what may proceed, what must be blocked, what requires validation, what can be reused, and what can become canonical.
Because promotion transfers authority, each level increases burden. A signal can be weak. A candidate must be clear. A supported constraint must have evidence. A validated constraint must have a test. An enforced constraint must have an enforcement point. A canonical constraint must hold across contexts, survive conflict, have independent support, and possess a lifecycle owner.
The promotion ladder remains:
| Level | Meaning |
|---|---|
| Signal | Observed pattern, failure, intuition, or candidate rule. |
| Candidate | Clear proposed constraint with source and scope. |
| Supported | Candidate with repeated or strong evidence. |
| Validated | Supported constraint tested through explicit validation. |
| Enforced | Validated constraint with operational consequence. |
| Canonical | Enforced constraint stable enough to become system authority. |
| Deprecated | Constraint loses authority due to invalidation, replacement, context change, or supersession. |
The promotion evaluation engine makes this operational. It checks statement clarity, source, evidence threshold, scope, conflicts, validation, enforcement, and canonicality. Its outputs include PROMOTE, PROMOTE_WITH_LIMITS, DEFER_PROMOTION, ESCALATE_PROMOTION, BLOCK_PROMOTION, and DEPRECATE. This matters because promotion should not be binary. A constraint may be valid within a limited scope without becoming canonical system-wide.
Promotion must block when a constraint is ambiguous, sourceless, untestable, unenforceable, or overbroad. It must defer when evidence, validation, or scope is incomplete. It must escalate when conflicts, high-impact consequences, ethical uncertainty, legal uncertainty, or irreversible effects require human or governance authority. It must deprecate when a constraint is contradicted, superseded, invalidated, or context-expired.
The key rule is that promotion increases burden, not confidence. The higher a constraint rises, the less ambiguity is tolerated. A constraint that depends only on model behavior cannot be canonical. A constraint with no enforcement point cannot be enforced. A constraint with hidden conflicts cannot be promoted safely. A constraint with unclear scope may still be useful, but it must remain bounded.
This promotion layer is how Substrate avoids two opposite failures. The first failure is canon inflation, where too many ideas become “canonical” because they sound important. The second is endless provisionality, where nothing gains authority because the system refuses to decide. Promotion solves both by defining what evidence is required for each level.
In the current stage, most new items should remain signals, candidates, supported constraints, or process-enforced rules. Very few should be treated as canonical. The purpose of the supplemental is therefore not to promote everything. It is to preserve candidates, identify what has already earned authority, and create a clean queue for future Canon v5 synthesis.
Part XI — Final Operational Compression
The supplemental execution layer can be compressed into one principle:
Substrate advances only when mechanism, proof, and enforcement remain aligned.
Mechanism without proof is demo theater. Proof without enforcement is advisory process. Enforcement without visibility is hidden control. Visibility without behavior change is governance theater. Expansion without causal signal is drift.
The current system must therefore preserve the following sequence:
build one slice → prove the slice → test excluded behavior → evaluate causal value → preserve failure → decide whether expansion is justified
This sequence is the operating bridge between Canon v4 and future Substrate. It explains why the system must remain strict, why proof must be comparative, why PRs must be isolated, why future systems must remain deferred, and why failure must become system input.
The practical rules are:
- Build only what the current stage requires.
- Treat every PR as a boundary event.
- Require proof before accepting capability.
- Require negative validation before trusting scope.
- Treat technical completion as prerequisite, not proof.
- Compare reuse against baseline before claiming value.
- Treat unclear results as failure for planning.
- Ground all reasoning in actual system state.
- Convert failures into constraints, validation, redesign, or stop conditions.
- Promote constraints only when they earn authority.
This is the discipline that allows Substrate to grow without losing interpretability.
Current Build Guidance
Current active concern: PR9.5–PR11 / Stage 0 proof.
Current mechanism under test: artifact → reuse → improved reasoning.
Do not build yet: full claims, full trace, decision events, federation, ingestion, advanced UI, accounts, payment.
Primary proof requirement: baseline vs reuse comparison.
Primary failure to avoid: structured clarity mistaken for improvement.
Current enforcement reality: mostly process + visibility, not full system enforcement.
Stop condition: if reuse is not causal, do not expand.
Addendum A — Constraint Extensions: Delta from Canon v4
Purpose
This addendum preserves the constraint material that is not fully contained in Canon v4 but is too important to lose when prior working documents are archived. These items should be treated as supplemental constraint candidates, not as a replacement for the canonical system laws. Their function is to record what execution, validation, PR review, and failure analysis have revealed since the earlier canon work.
Canon v4 defines the governing ontology: what a constraint is, how constraints are typed, how they are promoted, how validation works, and how enforcement grants consequence. This addendum records the operational constraints that emerged while testing whether those concepts survive implementation. Some of these may eventually be folded into Canon v5. Others may remain process doctrine or stage-specific execution rules.
The key distinction is that Canon v4 mostly defines system law, while this addendum preserves execution law: what must happen for those laws to remain true during build, review, testing, and expansion.
A.1 Constraint Extension Classification
The supplemental constraints fall into five groups.
First, there are mechanism constraints, which determine whether the core loop is real. These include causal reuse, visible reuse, artifact usefulness, and baseline comparison.
Second, there are execution constraints, which determine whether PRs and staged development preserve causal clarity. These include slice isolation, no scope contamination, proof over plausibility, and negative validation.
Third, there are epistemic constraints, which determine whether outputs, measurements, and validations can be trusted. These include measurement skepticism, caveated validation, agreement-is-not-validation, and coverage honesty.
Fourth, there are expansion constraints, which determine when it is safe to increase system complexity. These include signal preservation under expansion, no future-based justification, and elimination over expansion.
Fifth, there are human/system interaction constraints, which determine whether humans, interfaces, and Aalam behavior preserve or degrade system reliability. These include visibility defines reality, manual over automated, structured human oversight, and weak-HITL awareness.
A.2 Supplemental Constraint List
SC-01 — Structure Must Do Work
A structure is valid only if it changes system behavior, enables validation, constrains output, improves reuse, or reduces failure. A schema, label, checklist, artifact type, or issue category that merely organizes text is not yet a governed structure.
This constraint preserves a central lesson from the MVP and canon work: Substrate is vulnerable to fake rigor. The system can produce structured prose, schemas, issue grids, roadmap maps, and evaluation tables that appear disciplined but do not alter behavior. Structure becomes meaningful only when it creates a new capability, prevents a failure, or enables a test.
Status: supplemental candidate for Canon v5.
Primary source cluster: constraint consolidation and MVP execution docs.
Failure if missing: process theater, schema bloat, false confidence.
SC-02 — Proof Over Plausibility
No PR, artifact, system capability, validation result, or architectural claim should be accepted because it appears plausible. The proof must be explicit, reproducible, and tied to observable behavior. Valid proof includes database evidence, API responses, test output, screenshots where relevant, saved artifacts, or reproducible user-visible behavior. Invalid proof includes UI appearance alone, code presence alone, expected behavior, or developer confidence.
This began as an MVP PR rule, but it is broader than the MVP. It is a permanent system discipline because Substrate’s value depends on distinguishing what exists from what merely appears to exist.
Status: active execution constraint; candidate for Canon v5 integration.
Primary source cluster: execution governance and system-state docs.
Failure if missing: demo illusion, premature progression, hidden implementation gaps.
SC-03 — Negative Validation Is Required
Positive success does not prove a system slice. The system must also prove that out-of-scope behavior is inactive. If reuse is not in scope, reuse must not occur. If automation is not in scope, automation must not occur. If claim extraction is not active, outputs must not be treated as structured claims.
This constraint protects sequencing. Without negative validation, future logic can leak into current work and make later proof uninterpretable. A PR can pass its intended test while contaminating the system with hidden future behavior.
Status: active execution constraint; already implied by Canon v4 falsification, but stronger here.
Primary source cluster: PR4–7 execution plans and operating doctrine.
Failure if missing: scope contamination, hidden coupling, false causal attribution.
SC-04 — Reuse Must Be Causal
Artifact reuse is real only if the reused artifact materially changes output in a way that can be attributed to that artifact. Reuse must not be inferred from improved output alone. It requires a baseline condition, a reuse condition, visible artifact insertion, and a meaningful output delta.
This is one of the most important supplemental constraints because artifact reuse is the primitive on which the rest of Substrate depends. If reuse does not improve reasoning, preserve structure, reduce effort, or increase continuity, then artifacts may still be useful as notes, but they do not justify the broader system.
Status: active Stage 0 / PR10–PR11 gate; strong candidate for Canon v5.
Primary source cluster: MVP validation addendum, evaluation protocol, proof register.
Failure if missing: reuse illusion, placebo effect, structured inefficiency.
SC-05 — Visibility Defines Operational Reality
A capability that is not visible, user-triggerable, and inspectable should not be treated as active. The current system-state document states the operational version clearly: what is not implemented and user-visible does not exist. This applies especially to artifact creation, artifact retrieval, reuse, validation status, decision state, and authority.
Visibility does not mean every internal mechanism must be exposed. It means system-relevant state cannot remain hidden where it affects user reasoning or governance. Hidden state undermines reproducibility and weakens validation.
Status: active system-state constraint; candidate for frontend/governance canon.
Primary source cluster: system-state grounding and MVP doctrine.
Failure if missing: hidden memory, implied authority, unverifiable reuse.
SC-06 — Slice Isolation Protects Causal Learning
Each PR must represent one minimal independent capability. It must not include partial future features, forward-looking logic that affects behavior, unrelated commits, or hidden dependencies. Slice isolation is not merely good engineering practice; it is how the project learns whether a mechanism works.
If slices are not isolated, success cannot be attributed. Later stages may appear to work because earlier stages quietly included part of them.
Status: active PR execution constraint.
Primary source cluster: execution governance docs.
Failure if missing: uninterpretable PR results, contaminated staging, false proof.
SC-07 — Technical Completion Is Not System Proof
A technically working system does not prove the system hypothesis. Chat, artifact creation, persistence, retrieval, and reuse must exist before evaluation begins, but their existence does not by itself validate Substrate. The proof condition is comparative: reuse must materially improve reasoning over baseline chat.
This constraint prevents the project from mistaking implementation completion for product or system validation.
Status: active evaluation constraint.
Primary source cluster: evaluation protocol and MVP validation addendum.
Failure if missing: feature completion mistaken for mechanism validation.
SC-08 — Caveated Validation Is Valid Only When the Caveat Is Bound
A result may be accepted as caveatedly valid only when the core mechanism is proven, the weakness is explicit, the weakness is bounded, and the limitation is assigned to follow-up. A caveat is invalid if it undermines the conclusion, is hidden, or allows a weaker result to be presented as stronger than it is.
This constraint creates a middle path between false binary thinking and false confidence. It allows progress without corrupting the proof record.
Status: active review constraint.
Primary source cluster: validation results and artifact rules.
Failure if missing: over-blocking, under-blocking, or hidden epistemic weakening.
SC-09 — Signal Must Be Preserved Under Expansion
Expansion is valid only if causal signal remains attributable. New features, users, domains, agents, ingestion pipelines, UI layers, or governance mechanisms must not obscure the reason the system works or fails. Complexity is not neutral; it can destroy interpretability.
This constraint is central to deciding when to move from one stage to the next. Expansion is allowed only when the system can still identify whether value comes from artifact reuse, structured variation, validation, domain logic, or another named mechanism.
Status: active roadmap constraint; strong candidate for Canon v5.
Primary source cluster: validation results, post-MVP action plan, pressure register.
Failure if missing: expansion before proof, complex system with no causal diagnosis.
SC-10 — Insight Requires Structured Variation, Not More Output
Insight generation does not come from simply asking for more responses. It emerges when structured alternatives are generated, compared, and synthesized. Variation is therefore an extension of artifact reuse, not a separate uncontrolled mode. It should be triggered only when single-path reasoning is insufficient, uncertainty is high, multiple explanations are plausible, or stakes justify deeper analysis.
This constraint is important because it defines how Substrate may eventually become more than a memory or decision-support system. It can become an insight-generation system only if variation remains governed.
Status: future-facing mechanism constraint; belongs in supplemental and future roadmap.
Primary source cluster: validation results and variation trigger conditions.
Failure if missing: variant spam, redundant analysis, false depth.
SC-11 — Agreement Is Not Validation
Consensus among models, agents, reviewers, or users does not establish truth. Agreement may reflect shared training data, shared assumptions, shared bias, prompt convergence, social pressure, or selection effects. Consensus may be useful as a coordination signal, but it cannot replace evidence, executable testing, adversarial validation, or authority review.
This was already implied in Canon v4 through validation typing, but the failure taxonomy makes it operationally important. Synthetic consensus is one of the ways AI systems create false certainty.
Status: supplemental epistemic constraint; likely Canon v5 candidate.
Failure if missing: model agreement mistaken for proof.
SC-12 — Measurement Cannot Be Trusted by Default
Benchmarks, evaluation scores, green CI, ratings, and apparent performance may fail to measure the capability that matters. Measurement can create mirage effects: the system appears to perform a task while merely satisfying the metric.
This constraint is especially relevant to PR evaluation, artifact quality testing, and later domain systems. The system must validate the validator where the validator has high authority.
Status: supplemental validation constraint.
Primary source cluster: failure taxonomy, evaluation protocol, coverage honesty artifacts.
Failure if missing: benchmark theater, green-but-weak validation, overconfident stage transitions.
SC-13 — Coverage Honesty Is Required
Green CI or reduced test coverage does not prove validated behavior if the coverage was narrowed to obtain passing results. A green result under reduced coverage may be acceptable only if the reduction is explicitly recorded and bounded.
This rule belongs in the supplement because it is a concrete operational example of caveated validation. It protects the proof record from being laundered by technical success.
Status: active PR review constraint.
Primary source cluster: artifact set and caveated validation materials.
Failure if missing: technically green but epistemically weaker PRs.
SC-14 — Backend and Frontend Confidence Must Be Separated
A strong backend result must not launder a weak frontend result, and a polished frontend must not launder weak backend proof. Backend and frontend readiness should be evaluated separately when their proof quality differs.
This matters because Substrate depends on user-visible behavior and backend reality. Either side can create false confidence if collapsed into a single readiness judgment.
Status: active PR review constraint.
Primary source cluster: artifact set and system-state docs.
Failure if missing: full-stack readiness overclaimed from partial proof.
SC-15 — Decision Short-Circuiting Is Valid
Evaluation should stop at the first decisive failure condition when that failure blocks the result. This avoids over-analysis and prevents strong secondary signals from compensating for a failed hard gate. For example, if a PR violates scope, additional strengths do not rescue it.
This rule supports non-compensatory constraints. It is valuable for PR review, artifact evaluation, and stage gates.
Status: active review heuristic; compatible with Canon v4 hard constraints.
Primary source cluster: artifact set and PR governance.
Failure if missing: overthinking, compensatory scoring, weak artifacts admitted despite fatal gaps.
SC-16 — Human Oversight Must Be Structured, Not Assumed
Human-in-the-loop is weak unless humans are given structured evidence, explicit authority boundaries, and meaningful review surfaces. The governance-problems document identifies superficial oversight, cognitive overload, time pressure, and unclear responsibility as recurring failures.
This constraint prevents the system from treating “human review” as a magic safety layer. Human authority must be designed, scoped, and supported.
Status: supplemental governance constraint; important for Stage 4+.
Failure if missing: rubber-stamp review, overtrust, authority confusion.
SC-17 — Manual Over Automated Until Causality Is Proven
During early stages, manual explicit action is preferred over automation. The MVP operating doctrine required user-selected artifacts, user-initiated reuse, and no auto-retrieval or auto-injection.
This rule is not anti-automation. It is pro-causality. Automation hides causal paths too early. Once a mechanism is proven, automation may be introduced under constraints.
Status: active early-stage constraint.
Failure if missing: hidden memory, invisible reuse, non-reproducible improvement.
SC-18 — No Future-Based Justification
Work cannot be justified by future architecture, future scalability, future modules, or future governance value. A build item is valid only if it supports the current stage’s proof requirement.
This rule is essential because Substrate’s future architecture is large and tempting. Without this constraint, future vision will repeatedly leak into current execution.
Status: active stage discipline constraint.
Primary source cluster: operating doctrine and architecture vision.
Failure if missing: premature architecture, over-expansion, loss of proof focus.
SC-19 — Elimination Over Expansion
When uncertainty exists, the correct default is to remove scope, simplify, or defer. Expansion is justified only by evidence. The operating doctrine frames success partly by what is excluded, not what is added.
This constraint should remain active because Substrate’s complexity can grow faster than validation capacity.
Status: active stage discipline constraint.
Failure if missing: feature sprawl, complexity before signal, overbuilt MVP.
SC-20 — Artifact Quality Requires Heuristics, Not Just Structure
Validation results showed that artifacts became substantially more useful when they included lightweight decision heuristics, not only structure. A checklist can organize reasoning; a heuristic can execute judgment by identifying decisive failure conditions, priority order, or default assumptions.
This is a major insight for artifact design. Artifacts should preserve how to think, not just what was concluded.
Status: supplemental artifact-design constraint.
Failure if missing: artifacts reusable as templates but weak as decision tools.
A.3 Constraint Extensions That Belong in Future Doc 2
Some items from the uploaded materials are important but should not remain in the supplemental body because they define future systems rather than current or supplemental constraints. They should be preserved in the future work / capability roadmap document.
Items to move forward into Document 2:
- full claim schema
- full trace system
- decision-event schema
- full governed artifact lifecycle
- federation and institutional architecture
- detailed domain implementations
- ingestion pipelines
- advanced UI systems
- multi-agent governance
- Aalam variant architecture
- owner’s vault / historical corpus parsing
- payment and accounts
- personal / group / public spaces
These are deferred due to stage, not rejected.
Addendum B — Full Failure / Pressure Register
Purpose
This register preserves the failure and pressure material that should remain visible after earlier working docs are archived. It combines the full failure taxonomy with the operational failure/pressure register. Its purpose is not to catalog every possible problem; it is to identify the failures most likely to invalidate Substrate or distort its interpretation.
Every entry should be interpreted through four questions:
- What fails?
- How would the failure appear?
- What primitive or constraint responds to it?
- Does it block stage progression?
A future version of the register should be maintained as a live artifact.
B.1 Epistemic Failures
EF-01 — Hallucination
The model generates false information confidently. This is the baseline AI failure, but it is not the deepest Substrate problem. Hallucination matters because false outputs can become artifacts, reused context, or durable system memory if not constrained.
Observable symptom: confident unsupported statement.
Primitive response: claim isolation, evidence validation, artifact admission control.
Blocks progression: yes, if hallucinated claims enter reusable artifacts without status.
Source cluster: failure taxonomy.
EF-02 — Mirage Effect
The system appears to perform a task without actually doing it. Benchmarks, UI, or structured outputs may measure pattern recognition rather than real capability.
Observable symptom: system appears correct in demo or metric but fails under real task.
Primitive response: baseline comparison, adversarial validation, proof register.
Blocks progression: yes at PR9–PR11.
Source cluster: failure taxonomy and evaluation protocol.
EF-03 — Phantom Competence
High-quality outputs or high scores create an illusion of expertise. Even expert users can be misled when structure and fluency are strong.
Observable symptom: user treats structured output as more reliable than evidence supports.
Primitive response: epistemic state, uncertainty marking, evidence requirement.
Blocks progression: conditionally, especially in high-stakes domains.
Source cluster: failure taxonomy.
EF-04 — Epistemic Contamination Loop
AI hallucination becomes accepted by humans, enters documents or datasets, and is later reused as apparent knowledge. This is the “Bixonimania” pattern from the working docs.
Observable symptom: false information becomes durable source material.
Primitive response: artifact lineage, admission control, revocation, provenance.
Blocks progression: yes before ingestion or shared memory.
Source cluster: failure taxonomy and constraint consolidation.
EF-05 — Synthetic Consensus
Multiple models or agents produce the same answer, creating apparent validation. Agreement is mistaken for truth.
Observable symptom: consensus used as evidence.
Primitive response: evidence validation, adversarial validation, independent source check.
Blocks progression: yes for multi-agent validation if untreated.
Source cluster: failure taxonomy.
B.2 Reasoning and Cognitive Failures
RF-01 — Pattern Imitation Mistaken for Reasoning
LLMs reproduce plausible reasoning patterns without guaranteed underlying logic.
Observable symptom: coherent explanation fails under decomposition or counterexample.
Primitive response: decomposition, step validation, adversarial testing.
Blocks progression: yes for claim/reasoning layers.
Source cluster: failure taxonomy.
RF-02 — Distilled Error Propagation
Smaller or downstream models inherit hallucinations, biases, or reasoning errors from upstream models.
Observable symptom: repeated error across variants or agents.
Primitive response: independent validation, model role separation, provenance.
Blocks progression: later, especially with Aalam variants and docking.
Source cluster: failure taxonomy.
RF-03 — Overgeneralization / Trendslop
Outputs converge toward safe, generic, popular answers rather than context-sensitive reasoning.
Observable symptom: plausible but shallow answer; weak adaptation to context.
Primitive response: context binding, artifact-specific constraints, structured variation.
Blocks progression: conditionally, especially in domain systems.
Source cluster: failure taxonomy.
RF-04 — Anthropomorphic Projection
Users assume understanding, memory, intent, or authority where the system has none.
Observable symptom: users trust model behavior because it feels familiar or intelligent.
Primitive response: frontend visibility, explicit state, authority boundaries.
Blocks progression: yes before external use.
Source cluster: failure taxonomy and governance-problems doc.
B.3 Execution Failures
XF-01 — Fail-Open Execution
The system acts or proceeds despite missing validation, missing authority, or failed constraints.
Observable symptom: output, state change, or tool action occurs when preconditions are absent.
Primitive response: fail-closed gate, constraint evaluation engine, enforcement point.
Blocks progression: yes before any action-bearing stage.
Source cluster: failure taxonomy and evaluation engine.
XF-02 — Ambient Delegation Risk
Casual interaction creates real-world execution or implied authority without explicit decision.
Observable symptom: chat instruction becomes action path.
Primitive response: decision event, authority model, explicit approval.
Blocks progression: yes for tool/API stages.
Source cluster: failure taxonomy.
XF-03 — Valid Action, Wrong Outcome
A structured interface or API call is valid, but the action should not have occurred.
Observable symptom: syntactically correct operation produces inappropriate result.
Primitive response: pre-execution validation, post-execution verification, outcome checks.
Blocks progression: yes before external APIs and docking.
Source cluster: failure taxonomy.
XF-04 — Authority Escalation
Agents or users act beyond intended scope because boundaries are unclear.
Observable symptom: system performs or recommends action outside actor authority.
Primitive response: capability envelope, authority boundary, decision event.
Blocks progression: yes for governance stages.
Source cluster: failure taxonomy and constraint consolidation.
XF-05 — Hidden State Mutation
The system changes environment, data, memory, artifact state, or external system invisibly.
Observable symptom: state changes cannot be reconstructed or attributed.
Primitive response: trace, audit, immutable artifact versioning.
Blocks progression: yes before write actions or shared memory.
Source cluster: failure taxonomy.
B.4 System Design Failures
SF-01 — Tool Mistaken for System
An AI tool is deployed without process redesign, validation, artifact structure, or enforcement.
Observable symptom: useful outputs but no governed workflow.
Primitive response: docking harness, artifact layer, validation gates.
Blocks progression: yes before public positioning.
Source cluster: failure taxonomy and architecture vision.
SF-02 — Demo Illusion
The system works in ideal scenarios but fails under real use.
Observable symptom: polished demo, weak evaluation results.
Primitive response: real-task validation, baseline comparison, proof register.
Blocks progression: yes at Stage 0 exit.
Source cluster: failure taxonomy and evaluation docs.
SF-03 — Strategy Without Process Audit
Automation or system design is applied before understanding actual workflow.
Observable symptom: system optimizes wrong problem.
Primitive response: operating doctrine, real-use observation, staged roadmap.
Blocks progression: conditionally.
Source cluster: failure taxonomy and governance problem framing.
SF-04 — Discovery Without Constraint
Exploration proceeds without boundaries, causing chaotic behavior and scope drift.
Observable symptom: new primitives, features, or modules added before current proof.
Primitive response: single active objective, no future-based justification, elimination over expansion.
Blocks progression: yes if active during proof stages.
Source cluster: operating doctrine.
B.5 Organizational and Institutional Failures
OF-01 — Individual AI / Institutional AI Gap
Individuals become productive while organizations remain incoherent because artifacts, validation, and authority are not shared.
Observable symptom: private prompts improve local work but produce no institutional memory.
Primitive response: shared artifact standards, roles, workspaces, governance layer.
Blocks progression: later; relevant before institutional use.
Source cluster: failure taxonomy and architecture vision.
OF-02 — Signal Collapse / AI Slop
Output volume overwhelms evaluation capacity, causing signal scarcity.
Observable symptom: many outputs, few trusted artifacts.
Primitive response: artifact admission control, filtering, validation modes.
Blocks progression: yes before ingestion and public/shared systems.
Source cluster: failure taxonomy.
OF-03 — Sycophancy Bias Amplification
The system reinforces user beliefs or preferred narratives.
Observable symptom: answers increasingly agree with user framing without challenge.
Primitive response: adversarial review, Reviewer Aalam, contradiction tests.
Blocks progression: conditionally, especially in decision domains.
Source cluster: failure taxonomy.
OF-04 — Coordination Failure
Multiple agents, users, or groups act without shared standards.
Observable symptom: divergent artifacts, conflicting decisions, ungoverned parallel work.
Primitive response: role separation, shared constraints, consensus rules, federation later.
Blocks progression: later; especially group and institutional stages.
Source cluster: failure taxonomy and architecture vision.
OF-05 — Knowledge Fragmentation
Knowledge remains in private prompts, disconnected tools, or ephemeral chats.
Observable symptom: no shared institutional cognition.
Primitive response: artifacts, spaces, commonplace/public layer later.
Blocks progression: later, but motivates architecture.
Source cluster: failure taxonomy and architecture vision.
B.6 Human-Level Failures
HF-01 — Cognitive Debt
Users lose independent reasoning capacity through overreliance on AI.
Observable symptom: user accepts outputs without reconstruction or challenge.
Primitive response: structured interaction, artifact review, user checkpoints.
Blocks progression: long-term product risk.
Source cluster: failure taxonomy.
HF-02 — Orchestration Without Understanding
Humans manage agents or workflows without understanding the underlying reasoning.
Observable symptom: shallow control over complex system.
Primitive response: explicit artifacts, proof records, review surfaces.
Blocks progression: yes for multi-agent stages.
Source cluster: failure taxonomy and governance problem framing.
HF-03 — Weak HITL
Human-in-the-loop becomes superficial due to time pressure, overload, poor interfaces, or unclear responsibility.
Observable symptom: review exists but does not improve correctness.
Primitive response: structured authority, evidence presentation, decision events.
Blocks progression: yes before governance claims.
Source cluster: governance-problems doc.
HF-04 — Overtrust via Familiarity
The chat interface creates perceived safety and intimacy, causing risk to be underestimated.
Observable symptom: user treats system as safer or more knowledgeable than warranted.
Primitive response: UI state clarity, uncertainty, visible validation status.
Blocks progression: external testing risk.
Source cluster: failure taxonomy.
B.7 Market / Product Failures
MF-01 — Synthetic Value Illusion
AI-generated artifacts appear valuable because they are polished or abundant.
Observable symptom: output volume mistaken for product value.
Primitive response: artifact quality tests, reuse validation, user preference tests.
Blocks progression: yes before monetization.
Source cluster: failure taxonomy.
MF-02 — Oversupply / Commoditization
Massive output production reduces differentiation and drowns signal.
Observable symptom: many artifacts, few worth reusing.
Primitive response: filtering, promotion, artifact pruning, domain constraints.
Blocks progression: later.
Source cluster: failure taxonomy.
MF-03 — Value Distribution Skew
Infrastructure providers capture value while users produce content or reasoning artifacts.
Observable symptom: product value difficult to capture despite user activity.
Primitive response: product positioning, domain focus, institutional memory layer.
Blocks progression: business risk, not core system blocker.
Source cluster: failure taxonomy.
B.8 Governance Failures
GF-01 — Missing Validation Layer
Outputs proceed without systematic correctness checks.
Observable symptom: artifacts, claims, or decisions accepted without validation mode.
Primitive response: validation registry, proof register, evaluation engine.
Blocks progression: yes beyond early MVP.
Source cluster: failure taxonomy and evaluation engine.
GF-02 — Missing Decision Authority
The system cannot determine who decides, when, and under what authority.
Observable symptom: model output treated as decision.
Primitive response: decision event schema, authority boundary.
Blocks progression: yes before action or governance layers.
Source cluster: failure taxonomy and evaluation engine.
GF-03 — Missing Enforcement
Rules exist but do not alter behavior.
Observable symptom: constraint documented but not blocking, routing, or escalating.
Primitive response: enforcement location registry, backend gates, process gates.
Blocks progression: yes before governance claims.
Source cluster: failure taxonomy and constraint consolidation.
GF-04 — No Fail-Closed Behavior
Errors are logged but not blocked.
Observable symptom: failure visible only after execution.
Primitive response: fail-closed design, no validation → no execution.
Blocks progression: yes before external actions.
Source cluster: failure taxonomy and evaluation engine.
GF-05 — Trace Without Action
Logs exist but do not affect admissibility, authority, or reuse.
Observable symptom: system can audit but not govern.
Primitive response: trace-to-enforcement binding.
Blocks progression: later, especially trace/governance stages.
Source cluster: failure taxonomy.
B.9 Operational Pressure Register
The following items were originally recorded as FP-01 through FP-14 and should remain active. They are repeated here in compressed operational form so they are not lost.
| ID | Pressure | Severity | Stage Relevance | Blocks? |
|---|---|---|---|---|
| FP-01 | Weak artifact degradation | Fatal | PR9.5, PR10, PR11, Stage 1 | Yes |
| FP-02 | Minimal compliance / user bypass | Fatal | All user-facing stages | Conditional |
| FP-03 | Reuse does not improve outcomes | Fatal | PR10, PR11, Stage 0 exit | Yes |
| FP-04 | Provenance ignored | Fatal | Trace/governance stages | Later |
| FP-05 | Wrong-but-plausible retrieval | Serious | Retrieval/reuse expansion | Yes when active |
| FP-06 | Ambiguous artifact state | Serious | Status/governance stages | Yes when active |
| FP-07 | Corpus degradation | Serious | Volume, ingestion, shared spaces | Eventually |
| FP-08 | Cognitive overhead exceeds value | Serious | All stages | Conditional |
| FP-09 | Local correctness ≠ system correctness | Serious | Claims/governance/action | Later |
| FP-10 | Governance theater | Serious | Governance UI/authority | Yes when active |
| FP-11 | Premature schema path dependency | Moderate | Schema/future layers | Conditional |
| FP-12 | Narrow user fit | Moderate | Product/domain expansion | No, but constrains claims |
| FP-13 | Structured clarity mistaken for improvement | Serious | PR9+, validation | Yes |
| FP-14 | Premature external/institutional use | Moderate | External testing/orgs | Yes until Stage 0 passes |
Source cluster: failure/pressure register.
B.10 Failure Register Rule
A failure register entry should not remain descriptive. Each entry must eventually be linked to at least one of the following:
- artifact admission rule
- validation mode
- PR gate
- frontend visibility rule
- backend enforcement point
- human authority checkpoint
- future-system dependency
- stop condition
If a failure has no action path, it remains an observation, not a governance input.
Addendum C — Failure → Primitive Mapping
Purpose
The failure register identifies what can go wrong. This addendum identifies what the system must contain in response. It converts failure classes into required primitives, enforcement implications, and stage relevance. This mapping should be preserved because it explains why Substrate’s architecture exists. Without it, the system can appear over-structured. With it, the structure becomes intelligible: every primitive exists because an unconstrained failure mode required it.
This addendum should not be read as an instruction to implement every primitive immediately. Some primitives are active now, some are partial, and some belong to future stages. The purpose is to preserve the causal chain from failure to system design.
C.1 Primitive Set: Current and Deferred
The working documents collapse the larger primitive set into a compact core. The strongest consolidated set contains fourteen primitives: constraint artifact, constraint regime, constraint strength, feasibility boundary, lineage, evidence, validation, enforcement point, promotion, conflict, lifecycle, authority boundary, capability envelope, and context/identity constraint.
For current execution, these reduce further into a practical working set.
Active or near-active primitives
| Primitive | Current role |
|---|---|
| Artifact | Durable reusable output; active primitive. |
| Reuse | Core mechanism under proof; active primitive. |
| Validation | Required for proof; partially active through process. |
| Failure | Must be preserved; active as register and evaluation category. |
| Constraint | Canonical concept; partially active through process. |
| Evidence | Required for proof and validation; active mostly through review artifacts. |
| Scope | Required for PRs, artifacts, and constraints; active process primitive. |
Partial primitives
| Primitive | Current role |
|---|---|
| Claim | Conceptually important, not fully implemented. |
| Trace | Minimal source linkage only; full trace deferred. |
| Promotion | Conceptually defined; process-level only. |
| Authority | Human/process authority active; system authority deferred. |
| Enforcement | Process-enforced now; system enforcement deferred. |
Deferred primitives
| Primitive | Deferred to |
|---|---|
| Decision event | Future authority layer. |
| Capability envelope | Future execution/governance layer. |
| Federation compact | Future institutional/federation layer. |
| Full governed artifact lifecycle | Future post-MVP governance layer. |
| Full claim graph | Future reasoning structure layer. |
This distinction must be preserved. Treating deferred primitives as active would create false system state. Treating partial primitives as absent would lose the system’s direction.
C.2 Epistemic Failure → Claim, Evidence, Validation
Epistemic failures occur when the system produces, preserves, or propagates apparent knowledge without proof. The failures include hallucination, mirage effects, phantom competence, epistemic contamination, and synthetic consensus.
The required primitives are:
- claim isolation
- evidence binding
- validation mode
- epistemic state
- artifact admission control
- contradiction preservation
The enforcement implication is that no claim-bearing artifact should be promoted, reused as authority, or incorporated into durable shared state unless its claims are either validated, explicitly provisional, or marked as uncertain. This does not require full claim schema immediately. It does require that the system avoid treating narrative output as validated knowledge.
Stage relevance:
- PR9.5: distinguish artifact vs claim.
- PR10–11: prevent reuse from laundering unsupported claims.
- Stage 3+: implement claim storage and validation.
- Domain systems: essential for language feedback and legislative effect claims.
C.3 Contamination Failure → Artifact Admission, Lineage, Revocation
Contamination failures occur when weak or false outputs become durable memory. The core pattern is: output appears useful, enters artifact storage, gets reused, and becomes harder to challenge later.
The required primitives are:
- artifact admission rule
- artifact quality floor
- lineage
- validation status
- deprecation / revocation path
- corpus hygiene
The enforcement implication is that artifact creation should not be treated as neutral. Even if early artifacts remain lightweight, the system must eventually distinguish between draft, useful context, validated artifact, and canonical or governing artifact. The failure register identifies weak artifact degradation as fatal because reuse over weak artifacts collapses the system into stored chat fragments.
Stage relevance:
- PR9.5: artifact quality classification.
- Stage 1: artifact consistency and signal filtering.
- Stage 2+: reuse restrictions and artifact state.
- Ingestion stages: mandatory before bulk corpus parsing.
C.4 Reasoning Failure → Decomposition and Structured Artifacts
Reasoning failures occur when outputs are coherent but not reliable. Pattern imitation, trendslop, overgeneralization, and anthropomorphic projection create outputs that look like reasoning but fail under pressure.
The required primitives are:
- decomposition
- assumptions
- constraints
- decision logic
- reusable heuristics
- structured variation where needed
The enforcement implication is that artifact quality cannot be measured only by readability. A strong artifact should preserve how to think, not merely what was concluded. The validation materials show that artifacts became stronger when they contained lightweight decision heuristics.
Stage relevance:
- PR9.5: test whether artifacts decompose.
- Stage 1–2: artifact quality and reuse design.
- Stage 3: claim and reasoning structure.
- Domain systems: language feedback and legislative drafting both require granular reasoning units.
C.5 Execution Failure → Decision Events, Capability Envelope, Fail-Closed Gates
Execution failures occur when the system acts without sufficient validation or authority. These include fail-open execution, ambient delegation, valid-action-wrong-outcome, authority escalation, and hidden state mutation.
The required primitives are:
- decision event
- authority boundary
- capability envelope
- pre-execution validation
- post-execution verification
- fail-closed enforcement
The enforcement implication is that generation and action must remain separated. The system may propose, structure, or evaluate, but it must not execute or authorize by implication. The constraint evaluation engine formalizes this: actions, promotions, and system changes require authority, evidence, capability-envelope compliance, enforcement binding, and decision events.
Stage relevance:
- Current stage: avoid direct execution.
- Stage 4+: decision events become active.
- Stage 5+: validation gates and enforcement become system primitives.
- External APIs / docking: mandatory before write access or external action.
C.6 System Illusion Failure → Proof Register and Comparative Evaluation
System illusion failures occur when the system appears to function but its mechanism has not been proven. Demo illusion, technical completion without value, and structured clarity without epistemic improvement all belong here.
The required primitives are:
- proof register
- baseline condition
- reuse condition
- evaluation artifact
- pass/fail/unclear state
- stop condition
The enforcement implication is that a system cannot advance because it works technically. It advances only if its mechanism produces observable value. The evaluation protocol explicitly states that the loop must produce visible, repeatable reasoning value over baseline chat.
Stage relevance:
- Stage 0 exit: mandatory.
- PR10–11: central.
- External testing: required before claims of product value.
- Monetization: should not proceed on demo value alone.
C.7 Human Oversight Failure → Structured Authority and Review Surfaces
Human failures occur when review becomes superficial, overloaded, unstructured, or biased by system fluency. Weak HITL is a central governance problem: humans skim, approve, or accept outputs without meaningful validation.
The required primitives are:
- explicit authority role
- review checklist
- evidence presentation
- escalation rule
- human decision record
- refusal / defer states
The enforcement implication is that “human review” is not a sufficient safety claim. Review must be structured, scoped, and recorded. If human review does not change behavior, it is not governance.
Stage relevance:
- Current stage: owner/process discipline.
- Stage 4+: decision authority.
- Domain systems: teacher/supervisor/legal expert roles.
- Institutional layers: org admin and head office authority.
C.8 Organizational Failure → Shared Standards, Spaces, Federation
Organizational failures occur when individual AI use improves local productivity but produces no shared institutional cognition. Knowledge fragments across private prompts, tools, chats, and local memories.
The required primitives are:
- shared artifact standards
- workspace boundaries
- role separation
- group/shared/public spaces
- institutional artifact lifecycle
- federation compact
The enforcement implication is that scaling Substrate beyond individual use requires more than accounts or collaboration UI. It requires governance of what artifacts mean, who may promote them, how conflicts are resolved, and what becomes shared knowledge.
Stage relevance:
- Stage 3+: small shared workspaces.
- Stage 4+: role separation.
- Stage 5+: institutional versions.
- Stage 8/federation: cross-node governance.
C.9 Market / Signal Failure → Filtering and Promotion
Market and signal failures occur when output abundance reduces value. AI slop, synthetic value illusion, oversupply, and signal collapse make evaluation the bottleneck.
The required primitives are:
- artifact quality scoring
- signal filtering
- promotion thresholds
- deprecation
- pruning
- user preference evidence
The enforcement implication is that more artifacts do not necessarily create more value. Corpus growth must be constrained by admission, reuse, and validation. The system should prefer fewer high-signal artifacts to many plausible weak ones.
Stage relevance:
- Stage 1–2: signal filtering.
- Ingestion: mandatory before large document parsing.
- Domain systems: necessary for corpora and learner/project memory.
- Public/commonplace layers: mandatory before publication.
C.10 Governance Failure → Evaluation and Promotion Engines
Governance failures occur when rules exist but do not determine admissibility, authority, or execution. Missing validation, missing decision authority, missing enforcement, no fail-closed behavior, and trace without action are all governance failures.
The required primitives are:
- constraint evaluation engine
- promotion evaluation engine
- enforcement location registry
- decision event
- trace update
- fail-closed default
The evaluation engine asks whether a claim, artifact, action, or promotion is admissible; it returns ALLOW, ALLOW_WITH_CONSTRAINTS, DEFER, ESCALATE, or BLOCK. The promotion engine asks whether a constraint has earned greater authority; it returns PROMOTE, PROMOTE_WITH_LIMITS, DEFER_PROMOTION, ESCALATE_PROMOTION, BLOCK_PROMOTION, or DEPRECATE.
Stage relevance:
- Current stage: conceptual/process enforcement.
- Stage 4+: decision events and authority.
- Stage 5+: system enforcement.
- Stage 6–7+: governance engine and docking.
C.11 Full Mapping Table
| Failure class | Required primitive | Enforcement implication | Stage relevance |
|---|---|---|---|
| Hallucination | Claim + evidence | No unsupported claim as authority | PR9.5, Stage 3 |
| Mirage effect | Proof register | Demo not accepted without comparison | Stage 0 |
| Phantom competence | Epistemic state | Confidence not treated as proof | All |
| Epistemic contamination | Artifact admission + lineage | Weak artifacts cannot become durable authority | Stage 1+, ingestion |
| Synthetic consensus | Independent validation | Agreement not proof | Multi-agent/federation |
| Pattern imitation | Decomposition | Reasoning must be inspectable | PR9.5, Stage 3 |
| Fail-open execution | Fail-closed gate | Missing validation blocks action | Stage 4+ |
| Ambient delegation | Decision event | Chat cannot imply authorization | Stage 4+ |
| Tool ≠ system | Docking harness | Outputs forced through governed pathway | Stage 6+ |
| Demo illusion | Baseline comparison | Technical working state not enough | Stage 0 |
| Weak HITL | Structured authority | Review must be scoped and recorded | Stage 4+ |
| Signal collapse | Filtering + promotion | Corpus growth requires admission | Stage 1+, ingestion |
| Governance theater | Enforcement binding | Visibility must affect behavior | Stage 4+ |
| Trace without action | Trace-to-enforcement link | Logs alone do not govern | Stage 5+ |
| Capability-control mismatch | Full constraint stack | No scaling without governance | All stage exits |
Addendum D — Enforcement Location Registry
Purpose
Canon v4 states that enforcement defines governance. This addendum specifies where enforcement can actually live. It preserves an important operational insight: constraints do not become real merely because they are named. They become real only when an enforcement surface blocks, alters, routes, escalates, or records system behavior in response.
At the current stage, many constraints are process-enforced rather than system-enforced. This is acceptable if explicit. It is dangerous if mistaken for implemented governance.
D.1 Enforcement Surfaces
Frontend Enforcement
Frontend enforcement governs what the user can see, trigger, inspect, and misunderstand. It is not cosmetic. It determines whether artifact creation, artifact selection, reuse, validation status, and failure state are visible at the moment they matter.
Frontend enforcement is appropriate for:
- artifact visibility
- reuse visibility
- status display
- warning and failure states
- user-trigger boundaries
- state labels
- comparison surfaces
- external tester UX
Frontend enforcement fails when UI implies authority, hides uncertainty, conceals reuse, or presents draft/weak artifacts as validated.
Backend Enforcement
Backend enforcement governs data integrity, ownership, state transitions, and admissibility. It is the strongest enforcement surface for durable system behavior because it can prevent invalid transitions regardless of UI.
Backend enforcement is appropriate for:
- artifact ownership
- artifact immutability/versioning
- persistence
- retrieval permissions
- validation status transitions
- promotion transitions
- decision-event requirements
- enforcement gates
Backend enforcement fails when models or UI are allowed to create state that the backend does not validate.
System Flow Enforcement
System flow enforcement governs sequencing. It ensures that operations occur in the correct order and that later stages cannot assume future capabilities.
System flow enforcement is appropriate for:
- create → store → retrieve → reuse sequencing
- no reuse before retrieval
- no evaluation before full loop
- no claim promotion before artifact quality
- no execution before decision
- no federation before core governance
System flow enforcement fails when work is parallelized in ways that destroy causal interpretation.
Process Enforcement
Process enforcement governs PR review, issue sequencing, proof discipline, and stop conditions. At the current stage, many critical constraints live here.
Process enforcement is appropriate for:
- slice isolation
- proof over plausibility
- negative validation
- no scope contamination
- reproducibility
- caveated merge rules
- issue cleanup
- stage exits
Process enforcement fails when human discipline weakens or when process rules are not converted into system enforcement when scale increases.
Human / Authority Enforcement
Human enforcement governs ambiguity, domain judgment, high-impact decisions, and situations where the system cannot determine legitimacy.
Human enforcement is appropriate for:
- authority decisions
- domain-sensitive validation
- cultural or educational judgments
- legal drafting review
- conflict resolution
- irreversible or high-impact actions
- canon promotion
Human enforcement fails when review is unstructured, undocumented, or treated as automatic validation.
D.2 Enforcement Types
| Enforcement type | Meaning | Strength |
|---|---|---|
| Hard block | System cannot proceed | Strongest |
| Soft block | System warns or requires confirmation | Medium |
| Visibility only | System exposes state but does not block | Weak |
| Test-only | Failure caught only in test/CI | Weak-medium |
| Process-only | Human review controls behavior | Stage-limited |
| Advisory | Rule exists but has no consequence | Not governance |
A constraint that remains advisory should not be described as enforced.
D.3 Enforcement Registry Template
Each important constraint should eventually have an enforcement record.
enforcement_record:
constraint_id:
constraint_statement:
enforcement_surfaces:
frontend:
status: missing | visible | blocking | not_applicable
mechanism:
failure_if_missing:
backend:
status: missing | visible | blocking | not_applicable
mechanism:
failure_if_missing:
system_flow:
status: missing | visible | blocking | not_applicable
mechanism:
failure_if_missing:
process:
status: missing | visible | blocking | not_applicable
mechanism:
failure_if_missing:
human_authority:
status: missing | visible | blocking | not_applicable
mechanism:
failure_if_missing:
current_enforcement_level: advisory | process | partial_system | enforced
stage_relevance:
owner:
next_action:
D.4 Enforcement Map for Supplemental Constraints
| Constraint | FE | BE | Flow | Process | Human | Current level |
|---|---|---|---|---|---|---|
| SC-01 Structure must do work | partial | no | partial | yes | yes | process |
| SC-02 Proof over plausibility | no | partial via tests | partial | yes | yes | process |
| SC-03 Negative validation required | no | partial via tests | yes | yes | yes | process |
| SC-04 Reuse must be causal | yes | partial | yes | yes | yes | process/partial |
| SC-05 Visibility defines reality | yes | partial | yes | yes | yes | partial |
| SC-06 Slice isolation | no | no | yes | yes | yes | process |
| SC-07 Technical completion is not proof | no | no | yes | yes | yes | process |
| SC-08 Caveated validation | no | no | no | yes | yes | process |
| SC-09 Signal preserved under expansion | no | no | yes | yes | yes | process |
| SC-10 Insight requires structured variation | future | future | future | yes | yes | conceptual |
| SC-11 Agreement is not validation | future | future | future | yes | yes | conceptual |
| SC-12 Measurement not trusted by default | future | partial tests | yes | yes | yes | process |
| SC-13 Coverage honesty | no | CI | no | yes | yes | process/CI |
| SC-14 BE/FE confidence split | yes | yes | no | yes | yes | process |
| SC-15 Decision short-circuiting | no | future | yes | yes | yes | process |
| SC-16 Structured human oversight | future | future | future | yes | yes | process/conceptual |
| SC-17 Manual over automated | yes | partial | yes | yes | yes | partial |
| SC-18 No future-based justification | no | no | yes | yes | yes | process |
| SC-19 Elimination over expansion | no | no | yes | yes | yes | process |
| SC-20 Artifact quality requires heuristics | yes | future | yes | yes | yes | process |
D.5 Current Enforcement Reality
The current system should be described honestly as follows.
Artifacts, storage, retrieval, and reuse are active or near-active primitives. Some constraints are supported by UI behavior and backend structure, but most governance rules are still enforced through process, review, and human discipline. There is not yet a full constraint enforcement engine. There is not yet a decision-event system. There is not yet a full trace layer. There is not yet a governed artifact lifecycle.
This does not invalidate the project. It defines the stage. It means the system can continue only if process enforcement remains strict until system enforcement replaces it.
D.6 Enforcement Migration Path
The system should migrate enforcement in this order.
Stage 0–1: Process and visibility
At the current stage, proof discipline, PR governance, artifact visibility, and reuse visibility are the main enforcement surfaces. This is sufficient only for controlled build and testing.
Stage 2–3: Backend state and artifact quality
As reuse becomes stable, artifact state, ownership, quality, and retrieval behavior should move into backend enforcement. The system should begin preventing weak state transitions rather than merely warning about them.
Stage 4–5: Authority and decision enforcement
When the system begins influencing real decisions, decision events, authority boundaries, and capability envelopes must become system objects.
Stage 6–7: Docking and external enforcement
When other models, tools, APIs, or external systems dock, outputs must be forced through a harness. This is where enforcement becomes structural rather than procedural.
Stage 8+: Federation and institutional enforcement
When systems operate across domains, organizations, or public/commonplace layers, enforcement must include cross-node standards, conflict handling, role hierarchy, and governance compacts.
D.7 Enforcement Rule
A constraint should not be called enforced unless at least one enforcement surface changes behavior when the constraint is violated.
If violation produces only explanation, logging, or regret, the constraint is not enforced. It is advisory.
Addendum E — Proof / Validation Register
Purpose
This addendum preserves the proof register as an operational instrument. Its purpose is to prevent vague optimism, demo theater, retrospective reinterpretation, and stage advancement based on technical completion alone. The register is not a general evaluation philosophy. It is a proof mechanism for deciding whether Substrate’s current mechanism has earned the right to expand.
The critical distinction is that implementation proves only that a system can execute. It does not prove that the system matters. The proof register determines whether the mechanism produces value over baseline AI use. In the current stage, that mechanism is explicit artifact reuse.
The register should be treated as active until the Stage 0 loop is fully validated and then adapted for later stages. It should not be archived as historical material.
E.1 Core Proof Question
The current proof question is:
Does explicit creation, persistence, retrieval, and reuse of artifacts materially improve reasoning outcomes over baseline chat?
The answer must be based on observable comparison, not impression.
The loop being evaluated is:
chat → artifact → storage → retrieval → reuse
Evaluation begins only after the full loop exists. If a user cannot create an artifact from real chat output, persist it, retrieve it later, explicitly reuse it, and see the reuse context, evaluation must not begin.
E.2 Scoring Rule
Use a 0–2 scoring scale unless otherwise specified.
| Score | Meaning |
|---|---|
| 0 | Failed, absent, or not observable |
| 1 | Partial, ambiguous, weak, or dependent on interpretation |
| 2 | Clear, observable, reproducible enough for current stage |
A proof run should record raw notes and scores. The statement “looks good” is not a valid evaluation result.
E.3 Proof Register Items
PV-01 — End-to-End Loop Existence
Question: Can one real user complete the full loop without hidden steps, mock behavior, or manual reconstruction?
Observation hooks:
- user sends chat message
- user saves artifact from real chat output
- artifact persists after refresh or new session
- user retrieves artifact later
- user explicitly reuses artifact in new chat
Scoring:
| Component | Score |
|---|---|
| Chat works | 0 / 1 / 2 |
| Artifact creation works | 0 / 1 / 2 |
| Persistence works | 0 / 1 / 2 |
| Retrieval works | 0 / 1 / 2 |
| Reuse works | 0 / 1 / 2 |
Pass threshold: all five items must score 2.
Fail condition: any item scores 0.
Inconclusive condition: any item scores 1 and none score 0.
Source: Proof / Validation Register.
PV-02 — Reuse Visibility
Question: Is reuse explicit, visible, and user-controlled rather than hidden, automatic, or merely claimed?
Observation hooks:
- artifact identity visible at point of reuse
- user explicitly selects artifact
- chat surface shows artifact is in use
- resulting output can plausibly be linked to the reused artifact
Scoring:
| Component | Score |
|---|---|
| Artifact selection visible | 0 / 1 / 2 |
| Reuse entry point visible | 0 / 1 / 2 |
| User control clear | 0 / 1 / 2 |
| Output influence legible | 0 / 1 / 2 |
Pass threshold: all four items score at least 1 and total score is at least 7/8.
Fail condition: any hidden or implicit reuse path exists in the tested flow.
Inconclusive condition: visibility exists but output influence is ambiguous.
PV-03 — Cross-Session Persistence
Question: Does artifact value survive the session boundary rather than only functioning as same-session convenience?
Observation hooks:
- artifact still present after refresh or new session
- user can reopen artifact without reconstruction
- artifact remains understandable outside its original session
Scoring:
| Component | Score |
|---|---|
| Artifact survives session boundary | 0 / 1 / 2 |
| Artifact can be reopened reliably | 0 / 1 / 2 |
| Artifact remains interpretable | 0 / 1 / 2 |
Pass threshold: total score at least 5/6 and no zeroes.
Fail condition: artifact disappears, corrupts, or becomes unusable across sessions.
Inconclusive condition: artifact persists but requires substantial manual reconstruction.
PV-04 — Baseline vs Reuse Comparison
Question: Does artifact reuse produce materially better performance than baseline chat on the same or equivalent task?
Observation hooks:
- same task run in baseline mode
- same task run in reuse mode
- comparison recorded in plain language
- artifact influence identified
Scoring dimensions:
| Dimension | Score |
|---|---|
| Clarity | 0 / 1 / 2 |
| Faithfulness to prior reasoning or intent | 0 / 1 / 2 |
| Effort reduction | 0 / 1 / 2 |
| Continuity across session boundary | 0 / 1 / 2 |
Pass threshold: reuse condition must score at least 2 points higher than baseline and must improve at least two dimensions by at least 1 point each.
Fail condition: no meaningful difference, worse result, or baseline preferred.
Inconclusive condition: minor difference without clear user preference or observable gain.
Source: Evaluation Protocol and Proof Register.
PV-05 — User Preference Under Real Use
Question: Would the user choose reuse over restating the problem from scratch for this task type?
Observation hooks:
- explicit user preference statement
- voluntary reuse attempt
- lower restatement burden
- continued use after novelty
Scoring:
| Component | Score |
|---|---|
| Stated preference for reuse | 0 / 1 / 2 |
| Voluntary second reuse attempt | 0 / 1 / 2 |
| Lower restatement burden reported or observed | 0 / 1 / 2 |
Pass threshold: total score at least 5/6.
Fail condition: user prefers starting over or avoids reuse despite successful mechanics.
Inconclusive condition: user accepts reuse once but shows no repeat preference.
PV-06 — Artifact Usefulness Quality Floor
Question: Are artifacts substantively useful rather than merely stored outputs?
Observation hooks:
- artifact understandable outside original chat context
- artifact carries enough structure to support later reuse
- artifact is not merely a transcript fragment, decorative label, or polished paragraph
Scoring:
| Component | Score |
|---|---|
| Self-contained enough for later reading | 0 / 1 / 2 |
| Supports later task continuation | 0 / 1 / 2 |
| Better than raw chat fragment | 0 / 1 / 2 |
Pass threshold: total score at least 5/6.
Fail condition: artifacts function only as storage, not reuse-ready reasoning objects.
Inconclusive condition: mixed artifact quality across runs.
PV-07 — Evaluation Readiness Gate
Question: Is the system ready to claim proof at all?
All must be true:
- user can create artifact from real chat output
- artifact persists after refresh or session change
- user can retrieve artifact later
- user can explicitly reuse artifact in chat
- output visibly reflects reuse
If any condition is false, evaluation must not begin and the stage cannot pass.
E.4 Evaluation Record Template
Each proof run should generate an evaluation artifact.
evaluation_record:
run_id:
date:
user:
task_type:
artifact_used:
baseline_summary:
reuse_summary:
pv_scores:
PV-01:
PV-02:
PV-03:
PV-04:
PV-05:
PV-06:
comparison:
clarity:
faithfulness:
effort:
continuity:
final_judgment: pass | fail | inconclusive
notes:
required_follow_up:
This record prevents future Aalam instances from reconstructing decisions from memory or tone. It also preserves the difference between a technical success and a system-validating result.
E.5 Valid Evidence of Reuse
Valid reuse evidence requires:
- real task
- baseline comparison
- observable improvement
- meaningful artifact contribution
- low adaptation cost
- natural use
- recorded comparison
Invalid evidence includes:
- trivial demo scenario
- reuse that does not change outcome
- reuse that requires heavy rewriting
- reuse performed only because the test demands it
- subjective improvement without observable difference
- same-session convenience mistaken for durable reuse
The MVP validation addendum states that at least one strong example must satisfy real task, observable improvement, meaningful artifact contribution, low adaptation cost, and natural use.
E.6 Reuse Failure Modes
Structural Overfit
Artifact is too tightly fitted to the original problem and does not transfer cleanly.
Content Dominance Over Structure
Artifact preserves the answer but not the reusable reasoning pattern.
Forced Reuse
Reuse happens only because the system is being tested.
High Adaptation Cost
Reuse requires as much effort as starting from scratch.
Ambiguous Artifact Selection
The user cannot tell which artifact helps.
Weak Transformation Capability
Artifact can be reused but adapts awkwardly.
No Outcome Improvement
Reuse happens, but the output is not better.
Loop Break
Artifacts are created but not naturally revisited.
If multiple failure modes appear in a scenario, reuse is not validated. The correct response is not expansion; it is artifact redesign, interaction redesign, or reuse redesign.
E.7 Artifact Design Guidelines for Reuse
A good artifact should capture how to think, not merely what was concluded. It should preserve reusable reasoning structure, constraints, decision logic, and where possible, lightweight heuristics.
Each artifact should ideally include:
- problem type
- reusable structure
- constraints
- decision logic
- heuristic rule or priority order
- task-specific content separated from reusable structure
The artifact should pass this creation test:
Could this be reused tomorrow for a similar but different problem?
If the answer is no, the artifact is too vague, too specific, or too content-bound.
The validation materials showed that decision heuristics transformed artifacts from structural checklists into decision-executing frameworks.
Addendum F — PR Canon Checklist
Purpose
The PR Canon Checklist preserves the execution discipline required to prevent Substrate from becoming an unmanaged pile of partially connected work. A PR is not merely a code contribution; it is a stage-bounded system intervention. It must preserve causal clarity, prove its own slice, and avoid corrupting future interpretation.
This checklist should be used as the universal gate for PR work. It can be shortened for routine use, but the full version should remain preserved here.
F.1 Pre-PR Gate
Before an issue or PR is opened, the following must be true.
Scope
- The capability is required for the current stage.
- It does not depend on future-stage primitives.
- It can be described as one user-visible or system-visible improvement.
- It has a clear “does not do” boundary.
Stage Fit
- The work supports the current stage exit requirement.
- The work does not accelerate future architecture without proof.
- The work does not introduce dormant systems that affect behavior.
Proof Plan
- The proof method is known before implementation.
- At least one positive test is defined.
- At least one negative test is defined.
- Reproducibility is possible.
Failure Plan
- Likely failure modes are identified.
- The expected failure response is defined.
- If failure would block progression, that is stated.
If these are not true, the issue should be refined before implementation.
F.2 PR Body Requirements
Every PR should answer the following in the PR description.
## Scope
What this PR does:
What this PR explicitly does NOT do:
## Proof
What evidence shows the capability works?
## Negative Validation
What excluded behavior was tested and shown inactive?
## Reproducibility
How can another actor reproduce the result?
## Canon / Constraint Impact
Which constraints are preserved, tested, pressured, or changed?
## Failure Notes
What failed, what remains weak, and what is deferred?
This should be compact but mandatory for stage-relevant PRs.
F.3 Merge Gate
A PR can merge only if all are true:
- scope matches the intended slice
- proof is explicit
- negative test exists where relevant
- implementation matches intent
- no hidden or partial future feature affects behavior
- behavior is reproducible
- frontend and backend confidence are not collapsed
- caveats are explicit and bounded
A PR should be blocked if any hard gate fails. A PR may merge with caveat only if the core slice is proven and the caveat does not undermine the conclusion.
F.4 Universal PR Questions
Every PR must answer:
- What does it do?
- What does it not do?
- Where is the proof?
- What happens when it is broken?
- What future behavior must remain inactive?
- Does this create or alter a durable artifact?
- Does this affect authority, validation, or reuse?
- Does frontend state match backend reality?
- Does green CI actually cover the behavior being claimed?
- Does the PR preserve the current stage’s causal signal?
These are not bureaucratic questions. They are the mechanism that keeps system development interpretable.
F.5 PR Failure Patterns
UI Illusion
The UI appears to work, but persistence or backend state does not.
Response: reject or require backend proof.
Hidden Coupling
Future logic exists and may affect behavior even if not surfaced.
Response: reject or make inert.
Missing Negative Test
Only the happy path is tested.
Response: require excluded behavior test.
Dirty Scope
PR includes unrelated changes, future scaffolding, or mixed commits.
Response: split or reject.
Backend / Frontend Proof Mismatch
One layer is strong and the other weak.
Response: evaluate separately and caveat or block.
Green but Weak
CI passes because coverage is reduced or behavior is not actually covered.
Response: record caveat or require coverage restoration.
Over-Guidance
Developer is instructed step-by-step rather than given outcome, constraints, and proof criteria.
Response: return to outcome-based instruction.
The execution docs identify many of these patterns as central risks in PR4–7 and beyond.
F.6 Post-Merge Verification
After merge, at least one verification should occur.
- Confirm the capability still works in integrated environment.
- Confirm excluded behavior remains inactive.
- Confirm no regression in prior slices.
- Confirm artifact/state created by the PR is inspectable if relevant.
- Record caveats or follow-ups.
Post-merge verification is especially important when the merge was caveated.
F.7 Canon Impact Section
Every meaningful PR should include a short Canon Impact section.
## Canon Impact
This PR:
- [ ] preserves existing canon
- [ ] tests existing canon
- [ ] pressures existing canon
- [ ] requires canon revision
- [ ] introduces a new canon candidate
Constraints touched:
Evidence:
Open pressure:
The point is not to expand canon constantly. The point is to prevent unexamined canon drift.
F.8 PR Checklist Compression
For routine use, the checklist compresses to four questions:
- What does it do?
- What does it not do?
- Where is the proof?
- What happens when tested against its boundary?
If any answer is unclear, the PR is not ready.
Addendum G — Validation Mode Registry
Purpose
Canon v4 defines validation modes. This addendum preserves their operational use, limits, and failure modes. The registry exists because “validate this” is not a sufficient instruction. Validation must be typed by the object, risk level, context, and expected failure mode.
The key principle is that validation mode mismatch is itself a failure. A factual claim, a code path, a candidate selection, a legislative provision, a language correction, and a governance decision cannot all be validated the same way.
G.1 Evidence-Based Validation
Evidence-based validation is used for factual claims, source claims, beneficiary claims, language correction claims, legislative effect claims, or any assertion that can be supported or contradicted by evidence.
Requires:
- claim isolation
- source evidence
- evidence sufficiency
- contradiction check
- provenance
Strengths:
- appropriate for factual correctness
- supports claim-level accountability
- prevents unsupported narrative from becoming authority
Failure risks:
- claim is compound
- evidence supports only part of claim
- source is weak or misread
- absence is mistaken for contradiction
- retrieved evidence is incomplete
Best used when: the object is a factual or evidentiary assertion.
Weak when: the object is a plan, selection, design choice, or value judgment.
G.2 Executable Validation
Executable validation is used when correctness can be tested by execution, schema, program, calculation, rule engine, or formal check.
Requires:
- executable representation
- defined expected result
- test case
- reproducible output
- failure output
Strengths:
- strongest where applicable
- reduces dependence on model judgment
- supports backend/system enforcement
Failure risks:
- test covers wrong behavior
- edge cases omitted
- green CI with reduced coverage
- schema passes but outcome is wrong
- executable check becomes proxy for deeper correctness
Best used when: the object has formal or machine-checkable behavior.
Weak when: judgment, meaning, context, or legitimacy is central.
G.3 Comparative Validation
Comparative validation is used when choosing among candidates, outputs, plans, artifacts, models, or variants.
Requires:
- candidate set
- comparison criteria
- selection rationale
- rejection rationale
- selection-failure check
Strengths:
- useful for artifact quality
- useful for variant testing
- supports better-than-baseline evaluation
Failure risks:
- candidate set is poor
- best candidate never generated
- criteria are wrong
- evaluator prefers fluent output
- comparison produces ranking, not truth
Best used when: the task requires selection among alternatives.
Weak when: truth must be established independently.
G.4 Adversarial Validation
Adversarial validation is used to test whether a system, artifact, prompt, constraint, or workflow survives pressure.
Requires:
- attack class
- adversarial input
- expected safe behavior
- failure response
- record of result
Strengths:
- exposes hidden fragility
- essential for boundary and authority tests
- prevents overfitting to happy paths
Failure risks:
- adversarial tests are too narrow
- attacks are unrealistic
- test becomes performative
- passing known attacks creates false security
Best used when: the system will face ambiguity, misuse, conflict, or scale.
Weak when: used without clear expected safe behavior.
G.5 Proxy / Metric Validation
Proxy validation is used when direct truth is unavailable and a measurable proxy is substituted.
Requires:
- proxy declaration
- scope of metric
- known blind spots
- Goodhart risk assessment
- revalidation schedule
Strengths:
- useful when direct validation is impossible
- can support early product evaluation
- allows rough signal detection
Failure risks:
- metric gaming
- benchmark theater
- proxy drift
- aggregate score hides local failure
- measurement mistaken for truth
Best used when: no direct validation exists and uncertainty is explicit.
Weak when: proxy is presented as proof.
G.6 Human / Authority Validation
Human validation is used when domain judgment, legitimacy, culture, law, pedagogy, ethics, or irreversibility matter.
Requires:
- named authority or authority class
- decision basis
- scope of authority
- evidence presented to reviewer
- record of decision
Strengths:
- handles ambiguity and legitimacy
- necessary for high-impact or domain-sensitive contexts
- supports external authority boundaries
Failure risks:
- rubber-stamp review
- wrong authority
- hidden value judgment
- inconsistent reviewer standards
- cognitive overload
Best used when: automated validation is insufficient and legitimacy matters.
Weak when: invoked as vague “human oversight.”
G.7 Consensus Validation
Consensus validation is used when agreement among actors, validators, institutions, or nodes matters.
Requires:
- quorum or agreement rule
- participant identity
- conflict rule
- ordering rule
- escalation path
Strengths:
- useful for governance and federation
- supports distributed legitimacy
- helps coordinate institutional decisions
Failure risks:
- agreement mistaken for truth
- capture
- groupthink
- shared bias
- unresolved minority conflict
Best used when: legitimacy or coordination requires agreement.
Weak when: factual correctness is the goal.
G.8 Runtime Validation
Runtime validation is used after deployment or execution to detect drift, degradation, violations, or unexpected effects.
Requires:
- monitoring signal
- violation detector
- rollback or mitigation path
- revision trigger
- trace update
Strengths:
- detects real-world behavior
- supports long-term governance
- catches failures not visible in pre-test
Failure risks:
- monitoring too weak
- no rollback
- violations logged but not acted on
- signal ignored due to fatigue
Best used when: system operates over time or affects real workflows.
Weak when: no intervention path exists.
G.9 Validation Mode Selection Rule
Select validation mode by object type.
| Object | Primary mode | Secondary mode |
|---|---|---|
| Factual claim | Evidence | Adversarial / human |
| Code / schema | Executable | Adversarial |
| Artifact quality | Comparative | Evidence / adversarial |
| Reuse mechanism | Comparative | Runtime / human |
| Language feedback | Evidence / human | Runtime learner action |
| Legislative provision | Evidence / human | Adversarial / comparative |
| Promotion | Evidence / human | Cross-context / adversarial |
| External action | Executable / authority | Runtime |
| Federation decision | Consensus | Evidence / conflict review |
G.10 Validation Failure Rule
A validation result is invalid if:
- mode is undeclared
- object type does not match mode
- evidence is insufficient for risk level
- failure case is absent
- caveat undermines conclusion
- result cannot be reproduced
- validation does not affect decision
Validation that does not alter decision rights or system behavior remains advisory.
Substrate Constraint Canon v4 — Supplemental
Addenda Layer — Part 4
Addendum H: Failure / Repair Runs
Addendum I: Domain Constraint Extensions
Addendum J: Canon v5 Synthesis Queue
Addendum K: Deferred / Out-of-Scope Items for Document 2
Addendum H — Failure / Repair Runs
Purpose
This addendum preserves concrete examples of how Substrate should behave when the system fails or is repaired. The failure and repair runs are important because they show the difference between ordinary AI helpfulness and governed system behavior. A model can produce a useful answer while still failing governance. A governed system must expose what it knows, what it does not know, what it is allowed to do, and what must remain provisional.
Failure / repair runs should be retained because they make the canon operational. They demonstrate how constraints alter behavior.
H.1 Canonical Failure Run: Partial Source Treated as Full Authority
Scenario
A user refers to prior work, prior issue numbers, prior artifacts, or prior project structure. The assistant responds as if the full context is available, even though only partial context is visible in the current environment. It generates a confident issue definition, roadmap, constraint list, or artifact plan based on incomplete source visibility.
What Appears to Work
The output may be useful, coherent, and even directionally correct. It may follow known project patterns and produce a plausible artifact. It may satisfy the user’s immediate request. In ordinary AI interaction, this would often be treated as success.
What Fails
The system fails because it exceeds source authority. It converts partial visibility into apparent completeness. It treats inferred continuity as if it were source-grounded. It may collapse distinction between visible evidence, memory, inference, and reconstruction.
Failure Points
| Failure | Description | Required response |
|---|---|---|
| Source boundary failure | Visible context is incomplete but not declared. | Declare source limits. |
| Artifact authority failure | Generated artifact appears canonical without sufficient source. | Mark as draft or bounded reconstruction. |
| Trace failure | The path from source to statement cannot be reconstructed. | Require trace or scope limitation. |
| Validation failure | Claims are not tested against source material. | Mark unvalidated or defer. |
| Decision failure | Output implies action or acceptance without decision event. | Require explicit human/system decision. |
| Promotion failure | Draft artifact appears eligible for canon or issue creation. | Block promotion until source/evidence exists. |
Correct System Interpretation
This is not a failure to be helpful. It is a failure to preserve epistemic boundaries. The governed response should not refuse all work. It should continue within visible limits and label the output accurately.
H.2 Canonical Repair Run: Bounded Artifact Creation
Scenario
The same user request occurs, but the system recognizes partial visibility. Instead of pretending to possess full source authority, it creates a bounded artifact.
Correct Behavior
The system should state the visible basis, identify what is inferred, and separate confirmed structure from provisional reconstruction. It may still produce a useful issue, checklist, or document, but it must not present it as fully source-grounded unless it is.
Repair Pattern
| Step | Behavior |
|---|---|
| 1. Boundary declaration | State what source material is visible and what is not. |
| 2. Scope narrowing | Limit claims to visible evidence or explicitly marked inference. |
| 3. Artifact status | Label output as draft, reconstruction, or bounded recommendation. |
| 4. Validation path | Identify what would confirm or revise the artifact. |
| 5. Decision separation | Do not treat output as accepted, merged, or canonical without human decision. |
| 6. Failure preservation | Record missing context as a failure or caveat, not as invisible uncertainty. |
Correct Output Standard
A repaired output does not become weaker because it is caveated. It becomes more useful because its authority is honest. The system preserves progress without creating false certainty.
H.3 Failure Run: Reuse Illusion
Scenario
A user creates an artifact, retrieves it later, invokes reuse, and receives an improved output. However, there is no baseline comparison and no evidence that the artifact caused the improvement.
What Appears to Work
The interface works. The artifact exists. Reuse is visible. The resulting output is better than the prior output or feels more structured.
What Fails
The mechanism is unproven. The output may be better because of the model, because the user restated context, because the task is easier, because hidden context remains, or because the evaluation is subjective.
Required Repair
The repair requires a baseline condition and a reuse condition. The same or materially equivalent task must be run without the artifact and then with the artifact. The artifact’s contribution must be described. If the improvement is unclear, the result is inconclusive and must be treated operationally as failure for planning.
System Rule
Reuse is not proven by better output. Reuse is proven only by attributable delta.
H.4 Failure Run: UI Illusion
Scenario
The frontend shows an artifact, save action, status label, or reuse indicator, but backend state does not persist correctly or does not match what the user sees.
What Appears to Work
The user sees a functioning workflow. The UI gives confidence. The demo may look complete.
What Fails
The system does not actually preserve state. The artifact may be local, mocked, stale, incomplete, or disconnected from backend. This is especially dangerous because it creates false proof.
Required Repair
The repair requires backend proof: database record, API response, persisted ID, or cross-session retrieval. Frontend evidence is useful only when backed by persisted state.
System Rule
UI is not proof of system state. UI becomes proof only when it reflects backend truth.
H.5 Failure Run: Hidden Future Logic
Scenario
A PR for a narrow slice includes partial logic for a later stage. It may be described as inert, preparatory, or harmless. The future logic does not appear in the main user flow but affects models, schemas, endpoints, prompts, or state.
What Appears to Work
The PR may pass its immediate tests. The developer may believe future scaffolding saves time. The extra logic may seem harmless.
What Fails
Slice isolation fails. Later proof becomes contaminated because the future capability may no longer be introduced cleanly. The system loses the ability to know whether the current slice worked independently.
Required Repair
The repair is removal, isolation, or explicit inertness. If future fields exist, they must not affect behavior. If future endpoints exist, they must not be active. If future prompts exist, they must not influence output.
System Rule
Future logic that affects behavior invalidates the current slice.
H.6 Failure Run: Governance Theater
Scenario
The system displays badges, statuses, provenance labels, validation indicators, or approval language, but those surfaces do not alter behavior. Users begin treating the labels as meaningful.
What Appears to Work
The system looks governed. It shows structure, labels, and perhaps trace or validation status.
What Fails
The governance surface is descriptive rather than operative. If status does not change admissibility, reuse, promotion, or authority, it does not govern.
Required Repair
Bind status to behavior. If an artifact is unvalidated, it should behave differently from a validated artifact. If a claim is contested, it should not be reused as authority. If provenance is missing, the system should restrict promotion or require review.
System Rule
Governance surfaces must either constrain behavior or clearly state that they are informational only.
H.7 Repair Run: PR Evaluation Under Constraint
Scenario
A PR arrives that appears useful but contains partial future behavior, incomplete proof, or mismatched FE/BE confidence.
Correct Evaluation Sequence
- Identify the exact intended slice.
- Identify what the PR explicitly does not do.
- Verify proof for the intended slice.
- Run negative validation against excluded behavior.
- Separate backend proof from frontend proof.
- Check for hidden future logic.
- Record caveats if the core slice is proven but supporting proof is partial.
- Merge only if the caveat does not undermine the core conclusion.
Correct Output
The review should not say “good overall” or “looks fine.” It should classify the PR as mergeable, mergeable with caveat, blocked, or requiring split. The reasoning should state the decisive gate.
H.8 Repair Run: Artifact Quality Review
Scenario
An artifact is produced from chat and proposed for reuse. It is coherent and useful, but it may be too long, too content-bound, too specific, or too weakly structured.
Correct Evaluation
The artifact should be assessed for:
- problem type
- reusable structure
- constraints
- decision logic
- evidence or source basis
- applicability boundary
- minimal adaptation cost
- likely reuse scenario
Correct Result States
| Result | Meaning |
|---|---|
| Reject | Not useful as artifact. |
| Useful context | Useful to read but not reusable structure. |
| Reusable artifact | Can support future reasoning. |
| Claim-bearing artifact | Contains claims needing validation. |
| Validation artifact | Can be used to test or score. |
| Canon candidate | Strong enough to enter promotion path. |
System Rule
An artifact should not be promoted because it is well-written. It should be promoted only if it can guide future reasoning or constrain future behavior.
Addendum I — Domain Constraint Extensions
Purpose
This addendum preserves domain-specific constraint material without turning the supplemental into a domain implementation plan. Language acquisition and legislative drafting are important because they test whether Substrate’s general constraint system can specialize into real domains. They should not be built as full systems until the core loop and relevant stages justify them, but their constraints should be preserved because they reveal what the future system must support.
Domain constraints inherit the core canon. They cannot override core system laws. They can only specialize, refine, or strengthen them.
I.1 Language Acquisition Domain — Constraint Extensions
Domain Thesis
Language learning is controlled skill acquisition, not conversation. A language domain module must not be an “AI tutor persona” layered onto ordinary chat. It must be a governed learning loop: scenario, interaction contract, learner attempt, granular evaluation, feedback claim, scaffolded repair, language artifact, reuse, transfer test, validation, SkillGraph update, and progression.
This domain is valuable because it makes the core loop observable. Language learning naturally involves repetition, reuse, correction, transfer, and incremental difficulty. It is therefore a strong candidate domain after the structural core proves stable.
L-SC-01 — Every Language Interaction Requires a Contract
A language task must declare its mode, target skill, allowed input, expected output, evaluation method, and success condition. Without an interaction contract, evaluation becomes conversational and subjective.
Stage relevance: language module start.
System dependency: artifact types, validation modes, learner state.
L-SC-02 — Learner Attempts Are Artifacts
Every meaningful learner attempt should be preserved as a learning artifact. The original attempt, correction, context, target skill, validation state, and reuse history matter because they support evidence-based progression.
Stage relevance: language module MVP.
System dependency: artifact storage, learner memory, reuse.
L-SC-03 — Feedback Is a Claim
A correction is not merely a message. It asserts that something is wrong, why it is wrong, what rule or pattern applies, and what the learner should do next. Therefore, feedback requires confidence, evidence, and a validation task.
Stage relevance: language validation stage.
System dependency: claim-like feedback structure, human/authority fallback.
L-SC-04 — Feedback Must Be Validated Through Learner Action
A correction is not proven because the system gave it. It becomes useful only when the learner repairs, reuses, transfers, or stabilizes the form. This is a domain-specific version of causal reuse.
Stage relevance: language practice engine.
System dependency: learner attempts, repair loop, transfer tests.
L-SC-05 — Repair Must Be Bounded
The repair loop should move attempt → hint → retry → correction → validation. It must not allow infinite guessing or over-assistance. If the system immediately provides full answers, it may create correctness without learning.
Stage relevance: language module.
System dependency: scaffolding levels, attempt tracking.
L-SC-06 — Learner State Must Be Evidence-Based
A learner does not know a skill because they saw it once. Mastery requires evidence across recognition, guided production, independent production, transfer, and retention.
Stage relevance: learner memory stage.
System dependency: learner profile, history, review queue.
L-SC-07 — Evaluation Must Be Granular
Language output must be evaluated across layers: lexical, grammatical, syntactic, pragmatic, cultural, and pronunciation/phonological where relevant. A sentence is not simply right or wrong.
Stage relevance: language evaluation.
System dependency: granular feedback schema, conservative authority.
L-SC-08 — Transfer Must Be Tracked
Learning progression must test movement from repetition to substitution, recombination, cross-scenario use, and spontaneous use. This is where Substrate’s reuse primitive becomes educationally meaningful.
Stage relevance: i+1 practice and review.
System dependency: SkillGraph or lightweight skill dependencies.
L-SC-09 — Cognitive Load Must Be Constrained
Tasks must not introduce too much novelty at once. The system should control new vocabulary, new grammar, new context, and feedback density.
Stage relevance: language UX and practice design.
System dependency: interaction contract and learner state.
L-SC-10 — Authority Boundaries Must Be Explicit
AI may scaffold and propose language feedback, but culturally sensitive, dialectal, ambiguous, or low-confidence claims require uncertainty or human/community authority.
Stage relevance: public beta / real learners.
System dependency: confidence thresholds, human review.
I.2 Language Minimal Build Constraints
The language domain should begin with a small artifact system, not a full curriculum. The strongest early artifact types are:
- vocabulary artifact
- sentence-pattern artifact
- usage-contrast artifact
- learner-mistake artifact
The first language product should target intermediate learners, one language, small corpus, and artifact reuse. Full curriculum, audio, speech recognition, multilingual support, and heavy grammar engines should be deferred. The post-MVP plan warns that drifting into a full language app would weaken the Substrate-specific value.
Minimal loop:
ask → structured explanation → save artifact → reuse in new context → generate i+1 practice → track mistake / transfer
This belongs primarily in Document 2, but the constraint logic belongs here.
I.3 Legislative Drafting Domain — Constraint Extensions
Domain Thesis
Legislative drafting is not legal text generation. It is constrained institutional mechanism design. A statute, bill, amendment, or regulatory text does not matter only because it is well-written. It matters because it changes rights, duties, incentives, authorities, procedures, costs, enforcement conditions, and institutional behavior.
This domain is high-value because it pressures almost every Substrate primitive: claims, evidence, validation, trace, authority, conflict, enforcement, irreversibility, and public explanation. It should not be built before the relevant substrate layers exist, but its constraints should remain visible because they define a future high-stakes application.
LD-SC-01 — Draft Text Is Not Legislative Effect
The wording of a bill is not equivalent to what it will do. Legal text must be mapped to expected operational, fiscal, administrative, and social effects.
Stage relevance: legislative domain planning.
System dependency: claim/effect extraction, evidence validation, conflict analysis.
LD-SC-02 — Every Provision Requires an Effect Claim
Each provision should state what it intends to do, who it affects, through what mechanism, and under what conditions. Without effect claims, drafting remains text-level and cannot be validated.
Stage relevance: legislative claim layer.
System dependency: claim schema, effect taxonomy.
LD-SC-03 — Beneficiary Claims Require Validation
Claims about who benefits must be tested against the actual distribution of effects. Public descriptions often diverge from operational beneficiaries.
Stage relevance: legislative validation.
System dependency: evidence-based validation, data sources, human/legal review.
LD-SC-04 — Enforcement Must Be Explicit
A legal obligation without enforcement, remedy, penalty, funding, administrative capacity, or assigned authority may be symbolic rather than operational.
Stage relevance: legislative drafting module.
System dependency: enforcement analysis, authority mapping.
LD-SC-05 — Conflict With Existing Law Must Be Checked
Draft language must be tested against existing statutes, regulations, constitutional limits, administrative capacity, and judicial doctrine.
Stage relevance: legislative validation.
System dependency: external legal databases, trace, evidence, human authority.
LD-SC-06 — Incentive and Evasion Paths Must Be Modeled
Legislation must be tested for gaming, loopholes, burden shifting, unintended beneficiaries, and enforcement evasion.
Stage relevance: adversarial legislative validation.
System dependency: adversarial mode, structured variation, domain expertise.
LD-SC-07 — Source Influence Should Be Traceable
Lobbyist, donor, agency, interest group, model-generated, or constituent-originated language should be traceable where available. Source influence is not always invalid, but it must not be invisible.
Stage relevance: public/institutional legislative tools.
System dependency: provenance, source linkage, disclosure fields.
LD-SC-08 — Public Explanation Must Match Operative Text
The explanation of a bill must be validated against the actual legal mechanism. A public statement that claims one beneficiary, effect, or purpose while the operative text does something else is a governance failure.
Stage relevance: legislative transparency.
System dependency: claim/effect comparison, evidence binding.
LD-SC-09 — Irreversible or High-Impact Effects Require Stronger Validation
Criminal penalties, rights restrictions, taxation, surveillance, immigration, emergency powers, and irreversible administrative actions require elevated validation and explicit authority.
Stage relevance: high-impact legislative review.
System dependency: irreversibility gate, authority validation, human/legal review.
LD-SC-10 — Ambiguity Must Be Classified
Ambiguity may be intentional, unavoidable, harmful, or delegated. It must not remain invisible. Legislative ambiguity is not always a drafting defect, but it must be named.
Stage relevance: legislative drafting and review.
System dependency: ambiguity classification, human review.
I.4 Domain Integration Rule
Domain work should begin as diagnostic application, not full implementation. A domain can be used to test artifacts, reuse, validation, and constraint thinking before it becomes a product module. Full domain development should wait until the relevant substrate primitives exist.
Language can begin earlier because it uses the core loop directly and has lower institutional risk. Legislative drafting should begin as research, canon, and structured prototypes before it becomes an operational system, because it requires claim, trace, authority, evidence, and legal validation layers.
Addendum J — Canon v5 Synthesis Queue
Purpose
This addendum records what should be considered for eventual integration into Canon v5. It is not a mandate to rewrite Canon v4 immediately. Canon v4 should remain stable while PR9.5–PR11 and subsequent stages produce evidence. Canon v5 should be synthesized only when the supplemental constraints have been tested enough to distinguish stable laws from stage-specific doctrine.
J.1 Likely Canon v5 Additions
1. Structure Must Do Work
Canon v4 defines constraints and artifacts, but Canon v5 should make explicit that structure without behavioral consequence is not governed structure.
2. Reuse Must Be Causal
Artifact reuse should become a core system law or execution canon item. The system depends on reuse being attributable, visible, and materially useful.
3. Visibility Defines Operational Reality
Canon v5 should treat visibility as a governance requirement, not merely UI guidance. User-visible state, traceable reuse, and inspectable artifact status are conditions of proof.
4. Proof Over Plausibility
This may belong in the execution canon. It is currently process doctrine, but it functions as a system law during build and validation.
5. Negative Validation Required
Canon v4 already includes falsification and negative testing, but Canon v5 should strengthen it into a primary execution requirement.
6. Measurement Cannot Be Trusted by Default
Canon v5 should include a stronger warning that validation systems, benchmarks, green CI, and metrics themselves require validation.
7. Human Oversight Must Be Structured
Canon v5 should clarify that human-in-the-loop is not governance unless human authority is scoped, informed, and recorded.
8. Enforcement Location Must Be Declared
Canon v5 should add enforcement location as a required field or companion registry for every operational constraint.
J.2 Likely Canon v5 Refinements
Validation Modes
Canon v5 should preserve validation modes but clarify their hierarchy and misuse risks. Consensus validation should be explicitly marked weak for truth claims. Executable validation should be preferred only where the object is actually executable. Human validation should require authority definition.
Promotion
Canon v5 should integrate the promotion evaluation engine more directly. Promotion should remain transfer of authority, not recognition of usefulness.
Evaluation Engine
Canon v5 should preserve the evaluation engine but distinguish between current process-level evaluation and future system-enforced evaluation.
Domain Canons
Canon v5 should keep domain canons inherited, not parallel. Language and legislative constraints should specialize core laws rather than create separate governance systems.
J.3 What Should Not Be Added to Canon v5 Yet
The following should not be absorbed into Canon v5 until actual implementation pressure exists:
- full claim schema
- full trace schema
- decision-event schema details
- federation protocol
- detailed legislative drafting implementation
- detailed language product roadmap
- payment/accounts
- advanced UI states
- ingestion pipelines
- owner’s vault workflows
These belong in Document 2 as future systems and capability roadmap.
Addendum K — Deferred / Out-of-Scope Items for Document 2
Purpose
This addendum lists important material intentionally excluded from the supplemental body because it belongs in the future systems document. Exclusion here means “not current supplemental canon,” not “discarded.” These items should be carried into Document 2.
K.1 Full Claim Schema
The full claim schema should be preserved for future reasoning structure. It includes atomicity, source references, evidence anchors, epistemic state, scope, validation, trace reference, decision links, versioning, relationships, and integrity. It is not operationally active yet, but it will become relevant when artifacts need decomposition into claim-level units.
Document 2 placement: Stage 3 / claim layer, with minimal earlier use in PR9.5 artifact-vs-claim review.
K.2 Full Trace System
The full trace system should be preserved as future lineage infrastructure. It should eventually record source → transformation → artifact → claim → validation → decision → enforcement. It is not active now beyond minimal source references.
Document 2 placement: Stage 3–5, depending on when artifacts begin influencing decisions or shared memory.
K.3 Decision-Event System
Decision events should be preserved as the future authority layer. They become necessary when outputs become decisions, promotions, external actions, governance changes, or institutional commitments.
Document 2 placement: Stage 4 / governed execution.
K.4 Full Governed Artifact Lifecycle
The full lifecycle should be preserved for later: draft, admitted, validated, challenged, revised, superseded, deprecated, canonical. The current system should not implement this in full until artifacts have proven reuse value and state labels can alter behavior.
Document 2 placement: Stage 3–5.
K.5 Federation and Institutional Architecture
Federation, shared commons, cross-node validation, institutional memory, and public/commonplace layers should be preserved for later. They are not meaningful until the single-system governance loop is stable.
Document 2 placement: Stage 8+ / post-PR43.
K.6 Detailed Domain Implementations
Language and legislative domains should be preserved as future application tracks. Language may begin earlier as a low-risk domain test. Legislative drafting requires stronger substrate layers before operational use.
Document 2 placement: language from Stage 2/3 experiments to post-PR43 full domain; legislative as research/prototype earlier, operational later.
K.7 Ingestion Pipelines and Owner’s Vault
Historical project docs, owner’s vault, bulk ingestion, parsing, chunking, and extraction pipelines should be preserved but staged carefully. Ingestion before artifact quality and signal filtering risks corpus degradation.
Document 2 placement: light ingestion after signal quality; full ingestion after stronger artifact lifecycle and validation.
K.8 Advanced UI Systems
Advanced UI, dashboards, role-specific workspaces, group testing surfaces, public comparison UI, artifact graphs, and admin layers should be preserved as future product work. Early UI should remain simple and proof-oriented.
Document 2 placement: simple UI Stage 0/1; advanced UI Stage 4+ and public/product stages.
K.9 Accounts, Payment, and External Testing
Accounts, payment, group tests, public testing, and early monetization should be preserved for roadmap planning but should not be treated as constraint canon. They depend on system stability, evaluation evidence, and user value.
Document 2 placement: limited accounts/testing early if needed; payment after external value proof and reliability.
K.10 Multi-Aalam and Docking Harness
Aalam variants can begin as prompt-level experiments once evaluation exists, but governed variants and model docking require enforcement, trace, authority, and harness layers.
Document 2 placement: prompt-level variants Stage 2; governed variants Stage 4+; model docking Stage 4–6 depending on harness maturity.
Final Addenda Compression
The supplemental body explains the system. The addenda preserve the operating memory.
The required retained sets are:
- Constraint extensions
- Failure register
- Failure → primitive mapping
- Enforcement location registry
- Proof / validation register
- PR canon checklist
- Validation mode registry
- Failure / repair runs
- Domain constraint extensions
- Canon v5 synthesis queue
- Deferred future-system list
Together, these prevent the archived working documents from being lost. They also create the bridge from Canon v4 to future Canon v5 and from current Stage 0/7.1 execution to later architecture.
Member discussion: