1. Introduction

Substrate is a governed system for transforming language-based institutional artifacts—laws, policies, contracts, standards, investigative corpora, and AI governance rules—into explicit claims, constraints, provenance links, validation paths, failure tests, and uncertainty maps. Its purpose is not to decide cases, predict litigation outcomes, replace licensed experts, or determine political truth. Its purpose is to make the structure of consequential language visible, testable, and contestable before that language produces real-world harm.

The class of systems Substrate applies to can be summarized as:

language → interpretation → action → consequence

These systems include legislation, regulation, contracts, administrative rules, judicial interpretations, procurement frameworks, compliance policies, model governance specifications, large investigative datasets, scientific protocols, and institutional decision rules. They share a common failure pattern: language is treated as if it were executable logic, even though it is usually ambiguous, incomplete, context-dependent, incentive-sensitive, and unstable over time.

The core thesis is:

Institutional failure often arises not because rules are absent, but because rules remain implicit, ambiguous, unverifiable, unaudited, or misaligned with incentives. Substrate makes those hidden structures explicit without claiming final authority over their meaning.

Substrate begins modestly as an AI-assisted citizen drafting and review system: a user submits a proposed law or policy, and the system identifies undefined terms, internal contradictions, missing enforcement mechanisms, conflicts with existing law, interpretive uncertainty, and possible failure modes. Over time, the same architecture can grow into a federated legal-constraint network linking federal law, state law, bills, amendments, court interpretations, agency rules, budgets, contracts, enforcement data, international treaties, trade rules, and historical legal corpora.

The minimal product is already useful:

draft proposal
  ↓
claim extraction
  ↓
constraint check
  ↓
conflict / ambiguity / omission report
  ↓
rewrite suggestions

The long-term system is larger:

law-as-written
  + law-as-interpreted
  + law-as-funded
  + law-as-enforced
  + law-as-experienced
  + law-as-changed-over-time

Substrate is therefore not simply a legal technology project. It is a general architecture for systems where rules, language, incentives, and uncertainty interact under conditions of consequence.


2. Problem Definition

Language-based institutional systems fail in predictable, repeatable ways. These failures are not edge cases—they are structural.

2.1 Language Is Not Logic

Legal, policy, and institutional texts are written in natural language, which has properties incompatible with execution:

  • ambiguity (“reasonable,” “appropriate,” “necessary”)
  • context dependence (jurisdiction, time, precedent)
  • implicit assumptions (unstated definitions, norms)
  • layered meaning (text, intent, interpretation, enforcement)

Yet these texts are treated as if they define:

clear rules → consistent outcomes

In reality, they produce:

ambiguous rules → divergent interpretations → inconsistent outcomes

2.2 Interpretation Is Inevitable

No institutional system operates on text alone. Every rule is mediated by interpretation:

  • courts interpret statutes
  • agencies interpret regulations
  • companies interpret compliance rules
  • individuals interpret obligations

This creates a hidden layer:

text → interpretation → application

This layer is:

  • distributed (multiple actors)
  • inconsistent (jurisdictional variation)
  • dynamic (changes over time)
  • often undocumented in structured form

The result is that “what the law says” is not a single object—it is a space of interpretations.


2.3 Enforcement Is Incomplete

Even when rules are clear, enforcement introduces further gaps:

  • resource constraints (limited regulators, courts)
  • discretion (case-by-case decisions)
  • delay (litigation timelines)
  • selective enforcement (political or practical)

Thus:

rule ≠ enforced rule

2.4 Incentives Distort Outcomes

Actors respond to incentives, not just rules:

  • firms minimize compliance cost
  • agencies minimize workload or risk
  • legislators respond to donors, voters, or coalitions
  • litigants exploit ambiguity

This produces:

  • Goodhart effects (optimize the metric, not the goal)
  • strategic behavior (gaming definitions, thresholds)
  • omission (avoiding triggering conditions)
  • reinterpretation (narrowing or expanding scope)

Thus:

stated rule ≠ actual behavior

2.5 Drift Over Time

Rules and interpretations evolve:

  • amendments change statutes
  • courts reinterpret language
  • agencies update guidance
  • economic and technological contexts shift

This creates:

t₀: rule A
t₁: rule A'
t₂: rule A''

Often without explicit mapping between versions.

Drift can be:

  • explicit (amendments)
  • implicit (interpretation changes)
  • silent (appropriations, enforcement shifts)

2.6 Fragmentation Across Systems

Rules are distributed across:

  • statutes
  • regulations
  • judicial opinions
  • agency guidance
  • contracts
  • international agreements

These layers are:

  • partially overlapping
  • inconsistently defined
  • difficult to reconcile

Result:

multiple rule sources → unresolved conflicts

2.7 Lack of Structural Visibility

Current systems expose:

  • text (documents)
  • search (keywords)
  • summaries (LLMs)

They do not expose:

  • explicit claims
  • dependency structures
  • constraint relationships
  • conflict graphs
  • omission zones
  • uncertainty regions

Thus, users cannot easily answer:

  • What does this rule depend on?
  • Where does it conflict?
  • What is undefined?
  • How is it interpreted differently?
  • Where might it fail?

2.8 Consequence

The combined effect:

ambiguous language
+ distributed interpretation
+ incomplete enforcement
+ misaligned incentives
+ temporal drift
+ fragmented sources
--------------------------------
systemic uncertainty and failure

This leads to:

  • litigation as a primary resolution mechanism
  • inconsistent compliance outcomes
  • policy outcomes diverging from intent
  • opaque governance structures
  • high cost of understanding and participation

2.9 Core Problem Statement

The problem Substrate addresses is:

Institutional systems rely on language that is treated as if it were executable logic, but lacks explicit structure, validation, and failure detection—leading to hidden contradictions, inconsistent interpretations, and unpredictable outcomes.

Substrate does not eliminate ambiguity. It makes ambiguity explicit, structured, and testable.


3. Existing Systems and Limitations

A number of systems, standards, and platforms address parts of the problem Substrate targets. None integrate structure, constraints, validation, interpretation, incentives, and outcomes into a unified, governed system.

This section evaluates representative systems across four categories:

  • document standards
  • rule engines
  • data access platforms
  • governance / federation models

3.1 Akoma Ntoso / LegalDocML

What it does

Akoma Ntoso is an XML-based standard for structuring legal and legislative documents. It enables:

  • consistent tagging of sections, articles, clauses
  • metadata (authors, dates, jurisdictions)
  • cross-references within and across documents
  • machine-readable legislative corpora

Example:

<article id="art_1">
  <paragraph>
    <content>Employers must provide reasonable accommodations...</content>
  </paragraph>
</article>

What it does not do

  • does not extract or define atomic claims
  • does not encode constraints or logic
  • does not detect ambiguity or contradictions
  • does not represent interpretation or enforcement
  • does not support failure testing

Why it is insufficient

Akoma Ntoso structures documents, not meaning.

It answers:

“Where is this text located?”

It does not answer:

“What does this text do, and where does it fail?”

3.2 LegalRuleML

What it does

LegalRuleML provides a framework for representing legal rules and arguments in a formalized, machine-readable way.

It can express:

  • obligations, permissions, prohibitions
  • defeasible rules (rules with exceptions)
  • argument structures
  • logical relationships

What it does not do

  • does not originate from raw legislative text reliably
  • does not include a failure harness
  • does not integrate provenance, drift, or incentives
  • does not connect to real-world enforcement or outcomes
  • does not scale across heterogeneous domains

Why it is insufficient

LegalRuleML formalizes rules but assumes:

correct input → meaningful output

Substrate addresses:

ambiguous input → structured claims → tested constraints → surfaced failures

3.3 OpenFisca (Rules as Code)

What it does

OpenFisca converts tax and benefit legislation into executable code:

  • parameters (rates, thresholds)
  • eligibility rules
  • simulations of policy impact
  • API-based access

Example:

if income < threshold:
    benefit = rate * income

What it does not do

  • does not handle ambiguous or undefined terms
  • assumes rules are already formalized
  • does not detect contradictions in source law
  • does not model interpretation differences
  • does not expose incentive or governance layers

Why it is insufficient

OpenFisca operates on:

well-defined, parameterized rules

Substrate operates on:

messy, ambiguous, contested language before formalization

3.4 CourtListener / GovInfo / Congress APIs

What they do

These systems provide access to:

  • judicial opinions (CourtListener)
  • legislative texts and records (GovInfo, Congress APIs)
  • bill status, sponsors, votes
  • regulatory documents

They enable:

  • search
  • retrieval
  • citation tracking

What they do not do

  • do not extract claims or constraints
  • do not map dependencies across documents
  • do not detect conflicts or omissions
  • do not represent interpretation divergence structurally
  • do not connect law → incentives → outcomes

Why they are insufficient

They provide data access, not system understanding.

They answer:

“Where is the information?”

They do not answer:

“What does the system do, and where does it break?”

3.5 Estonia X-Road (Federation Model)

What it does

X-Road (X-tee) is a secure data exchange layer enabling:

  • interoperability across government systems
  • authentication and authorization
  • encrypted data transfer
  • logging and traceability
  • federation across institutions and countries

What it does not do

  • does not structure legal rules or claims
  • does not detect contradictions
  • does not model interpretation or incentives
  • does not provide validation or failure analysis

Why it is relevant

X-Road solves:

how systems connect

Substrate must solve:

what those systems mean and how they interact

X-Road is a transport layer, not a semantic or constraint layer.


3.6 Better Rules (New Zealand)

What it does

The Better Rules initiative focuses on:

  • making legislation understandable
  • translating policy into:
    • concept models
    • decision models
    • rule statements

It aims to make rules “digital-ready.”


What it does not do

  • does not include adversarial testing
  • does not detect Goodhart effects
  • does not map interpretation divergence
  • does not integrate multi-layer governance or outcomes
  • does not operate as a continuous failure-detection system

Why it is insufficient

Better Rules improves drafting clarity.

Substrate adds:

  • constraint enforcement
  • failure detection
  • cross-system linking
  • interpretation mapping

3.7 Summary Table

SystemStrengthMissing
Akoma Ntosodocument structuremeaning, constraints, validation
LegalRuleMLrule formalizationextraction, testing, integration
OpenFiscaexecutable rulesambiguity, interpretation, failure
CourtListener / GovInfodata accessstructure, analysis, linking
X-Roadsystem federationsemantic understanding
Better Rulesdrafting clarityadversarial robustness

3.8 Core Gap

Across all systems:

structure exists
rules exist
data exists

What does not exist is:

a system that:
- extracts claims from language
- builds constraint structures
- validates consistency
- detects failure modes
- maps interpretation differences
- links to incentives and outcomes
- preserves provenance and uncertainty

3.9 Positioning Substrate

Substrate is not:

  • a document standard
  • a rule engine
  • a data API
  • a legal search tool
  • a drafting assistant

It is:

a constraint-oriented system that sits above these layers and integrates them into a testable, failure-aware structure.

This distinction defines its role:

existing systems → provide inputs
Substrate → makes those inputs coherent, testable, and interrogable

4. Core Insight

Substrate is built on a single structural observation:

Many high-stakes systems rely on language that is treated as if it were executable logic, but is not.

This mismatch—between language and logic—is the root cause of the failures described in Section 2.


4.1 Language vs Structure vs Execution

It is useful to distinguish three layers:

Layer 1: Language
  - natural language text
  - ambiguous, flexible, context-dependent

Layer 2: Structure
  - claims, definitions, relationships
  - explicit but still non-executable

Layer 3: Execution
  - constraints, rules, enforcement
  - deterministic or semi-deterministic behavior

Most systems operate primarily at Layer 1. Some (e.g., LegalRuleML, OpenFisca) operate at Layer 3 but assume Layer 2 is already resolved.

The gap is:

Language → (missing transformation) → Structure → Execution

Substrate fills that transformation.


4.2 Claims as the Atomic Unit

Substrate treats claims as the smallest meaningful unit of analysis.

A claim is:

(subject, action, condition, exception)

Example:

Employer must provide accommodation unless undue hardship exists

Becomes:

subject: employer
action: provide accommodation
condition: employee has qualifying limitation
exception: undue hardship

This decomposition is critical because:

  • ambiguity becomes localized
  • dependencies become explicit
  • constraints can be constructed
  • validation becomes possible

4.3 Constraints as Executable Boundaries

A constraint defines:

what must hold
when it must hold
under what conditions it can fail

Example:

IF employee qualifies AND accommodation exists
THEN employer must provide accommodation
UNLESS hardship threshold exceeded

Constraints are not just rules; they are testable structures.


4.4 Validation vs Evaluation

Substrate distinguishes:

Evaluation: does this look correct?
Validation: can this be shown to hold under defined conditions?

Most current systems perform evaluation:

  • LLM outputs
  • summaries
  • heuristics

Substrate performs validation:

  • contradiction detection
  • missing definition detection
  • dependency resolution
  • failure simulation

4.5 Failure as a First-Class Object

Traditional systems treat failure as an exception.

Substrate treats failure as a primary output:

ambiguity
contradiction
omission
incentive misalignment
drift
governance gaps

The system is designed to answer:

Where does this break?

Not:

Is this good?

4.6 Interpretation as a Structured Layer

Interpretation is not noise—it is a necessary layer:

text → interpretation → application

Substrate captures interpretation as:

  • mappings from claims to interpretations
  • jurisdiction-specific overlays
  • temporal changes
  • divergence across authorities

This produces:

an interpretation space, not a single meaning

4.7 Incentives as Constraints

Rules do not operate in isolation. They interact with incentives:

rule + incentive → behavior

Substrate incorporates:

  • Goodhart effects (metric gaming)
  • strategic behavior
  • omission incentives
  • enforcement avoidance

Thus:

constraints must be evaluated in the presence of incentives

4.8 Provenance as a Requirement

Every claim, constraint, and interpretation must be traceable:

claim → source → version → context

Without provenance:

  • validation is impossible
  • disputes cannot be grounded
  • trust cannot be established

4.9 Uncertainty as an Output

Substrate does not aim to eliminate uncertainty.

It aims to structure it:

- undefined terms
- conflicting interpretations
- incomplete dependencies
- missing data

This produces:

uncertainty maps

These maps are more valuable than forced certainty.


4.10 Core Insight (Compressed)

Language-based systems fail because:
- meaning is implicit
- structure is hidden
- constraints are unenforced
- validation is absent
- failure is ignored

Substrate makes:
- meaning explicit
- structure visible
- constraints testable
- validation systematic
- failure central

This insight defines everything that follows.


5. System Overview (Substrate)

Substrate is a constraint-oriented compilation, validation, and failure-detection system that operates on language-based artifacts and produces structured, testable representations with explicit uncertainty.

At a high level, Substrate transforms:

raw language → structured claims → constraint system → validation → failure report → rewrite

5.1 System Definition

Substrate is composed of four core functional identities:

1. Compilation system
2. Constraint system
3. Validation system
4. Failure-detection system

Each operates on the output of the previous layer.


5.2 End-to-End Pipeline

Input (law, policy, contract, corpus)
  ↓
Segmentation
  ↓
Claim Extraction
  ↓
Type Assignment
  ↓
Constraint Construction
  ↓
Provenance Attachment
  ↓
Validation Engine
  ↓
Failure Harness
  ↓
Rewrite Engine
  ↓
Audit + Trace Output

5.3 Input Layer

Artifacts include:

  • statutes
  • regulations
  • contracts
  • policy documents
  • judicial opinions (later phase)
  • large document corpora (e.g., investigative datasets)

Requirements:

- ingest without loss
- preserve original structure
- maintain referenceability (sections, clauses)

5.4 Claim Compilation

The system extracts atomic units:

C1: issuer must provide coverage
C2: applies to group plans
C3: exception exists

Key properties:

  • decomposition into minimal units
  • explicit subject/action/condition/exception
  • ambiguity flagged, not resolved silently

5.5 Type System

Claims are typed:

- obligation
- permission
- prohibition
- definition
- exception
- dependency

This enables structured reasoning:

definition → used by obligation
exception → modifies obligation

5.6 Constraint Construction

Claims are assembled into executable logic:

IF condition
THEN obligation
UNLESS exception

This produces:

  • constraint graph
  • dependency graph
  • execution pathways

5.7 Provenance System

Every element is linked:

claim_id → source → section → version → timestamp

This ensures:

  • traceability
  • auditability
  • dispute grounding

5.8 Validation Engine

The system evaluates structural integrity:

- undefined terms
- missing dependencies
- circular logic
- conflicting constraints
- incomplete rules

Outputs:

validation_report:
  - errors
  - warnings
  - unresolved elements

5.9 Failure Harness

This is the defining component.

The system actively attempts to break the structure:

- ambiguity detection
- contradiction testing
- Goodhart simulation
- omission detection
- governance gaps
- drift sensitivity

Output:

failure_report:
  - failure type
  - location
  - impact
  - suggested fix

5.10 Rewrite Engine

Transforms failing structures:

ambiguous → parameterized
implicit → explicit
incomplete → extended

Produces:

  • structured rewrite (v0.2+)
  • diff from original
  • justification

5.11 Audit System

Captures full trace:

input → transformation → output → reason

Properties:

  • replayable
  • inspectable
  • immutable (ideally)

5.12 Governance Layer

Controls:

- who can modify
- who can override
- escalation rules
- human review points

5.13 Safety Core

Final enforcement:

block output if:
- missing fields
- unresolved conflicts
- incomplete validation

5.14 System Output

Instead of a binary judgment, Substrate produces:

- structured claims
- constraint graph
- validation report
- failure report
- rewritten version
- full audit trace

5.15 Key Property

Substrate ensures:

no output without:
- structure
- constraints
- validation
- failure analysis
- provenance

5.16 Minimal vs Full System

Minimal:

text → claims → basic validation → failure report

Full:

multi-layer system:
- law
- interpretation
- incentives
- outcomes
- drift
- federation

5.17 Positioning

Substrate sits above existing systems:

data sources → Substrate → structured, testable system view

It is not a replacement for:

  • legal systems
  • courts
  • experts

It is an infrastructure for:

making those systems structurally visible and interrogable.

6. Canon (Compressed)

The canon defines the non-negotiable system rules that govern Substrate. It is not a feature list; it is a constraint set that determines what the system must and must not do.

The full canon is provided in Appendix A. This section presents a compressed, structured version.


6.1 Canon Structure

The canon can be grouped into five domains:

1. Claims
2. Constraints
3. Validation
4. Audit
5. Governance

Each domain defines a class of system requirements.


6.2 Claims (Atomicity and Structure)

Core principles:

- All language must be decomposed into atomic claims
- Claims must be explicit, typed, and minimal
- No implicit assumptions are allowed to persist unexamined

Implications:

  • a sentence may contain multiple claims
  • undefined terms must be surfaced
  • claims must include:
    • subject
    • action
    • condition
    • exception (if present)

Failure condition:

If a claim cannot be decomposed, it cannot be validated.

6.3 Constraints (Enforceability)

Core principles:

- A rule is not a constraint unless it can be tested
- Constraints must be executable or evaluable
- Exceptions must be explicit and bounded

Implications:

  • “reasonable,” “appropriate,” “necessary” are not constraints
  • thresholds, conditions, and relationships must be defined
  • dependency structures must be explicit

Failure condition:

If a constraint cannot be evaluated, it is non-operational.

6.4 Validation (System Integrity)

Core principles:

- Validation is distinct from evaluation
- The system must detect:
  - contradictions
  - missing definitions
  - circular dependencies
  - incomplete logic

Implications:

  • multiple validation paths must exist
  • negative testing is required
  • local validation overrides global assumptions

Failure condition:

If a system passes evaluation but fails validation, it is unreliable.

6.5 Audit (Traceability)

Core principles:

- Every output must be traceable
- Every transformation must be recorded
- Provenance is required for all claims

Implications:

  • no anonymous outputs
  • no unexplained transformations
  • full decision paths must be reconstructable

Failure condition:

If a result cannot be traced, it cannot be trusted.

6.6 Governance (Control and Responsibility)

Core principles:

- Authority must be explicit
- Decision paths must be attributable
- Human oversight must be structured, not assumed

Implications:

  • escalation paths must exist
  • override mechanisms must be visible
  • responsibility cannot be implicit

Failure condition:

If no actor is accountable, the system is incomplete.

6.7 Cross-Cutting Principles

These apply across all domains:


6.7.1 Failure-First Orientation

The system prioritizes identifying failure over confirming correctness.

6.7.2 Uncertainty Preservation

Uncertainty must be surfaced, not suppressed.

6.7.3 Non-Substitution of Expertise

The system does not replace human judgment or licensed professionals.

6.7.4 Structural Integrity Over Output Fluency

A correct structure with visible uncertainty is preferred over a fluent but unverifiable answer.

6.8 Canon Compression (Single Statement)

A system is only reliable if:
- its claims are explicit
- its constraints are enforceable
- its validation is systematic
- its outputs are auditable
- its decisions are attributable
- and its failures are visible

6.9 Role of the Canon

The canon functions as:

- a design constraint
- a validation standard
- a development filter
- a failure detection baseline

It determines:

  • what the system can output
  • what must be blocked
  • what must be flagged
  • what must remain unresolved

This canon governs all subsequent architecture and behavior.


7. Architecture

Substrate is implemented as a layered constraint architecture, where each layer enforces a specific class of invariants. The system is not a pipeline of convenience; it is a sequence of gates that progressively transform and constrain language into a testable system.


7.1 Architectural Principle

Each layer must:
- introduce structure
- enforce constraints
- reduce ambiguity
- increase testability

No layer is optional. Skipping a layer produces structural failure downstream.


7.2 Layer Overview

L0: Input / Artifact
L1: Claim System
L2: Provenance System
L3: Constraint System
L4: Validation System
L5: Failure System
L6: Audit System
L7: Attribution System
L8: Governance System
L9: Optimization Boundary
L10: Safety Core

7.3 L0 — Input / Artifact Layer

Purpose

  • Ingest raw artifacts without loss

Inputs

  • statutes, regulations, contracts, policies
  • amendments, versions, metadata

Requirements

- preserve original text
- preserve structure (sections, clauses)
- maintain addressability

Failure condition

  • partial ingestion or loss of referenceability

7.4 L1 — Claim System

Purpose

  • Decompose text into atomic claims

Output

claim = {
  subject,
  action,
  condition,
  exception
}

Enforcement

  • reject non-decomposable constructs
  • flag ambiguity

Failure condition

  • unresolved composite statements

7.5 L2 — Provenance System

Purpose

  • Attach source and lineage to every claim

Output

claim_id → {
  source_document,
  section,
  version,
  timestamp
}

Enforcement

  • no claim without provenance

Failure condition

  • orphan claims or unverifiable sources

7.6 L3 — Constraint System

Purpose

  • Convert claims into executable or evaluable constraints

Structure

IF condition
THEN obligation
UNLESS exception

Output

  • constraint graph
  • dependency graph

Enforcement

  • explicit thresholds
  • explicit dependencies

Failure condition

  • non-operational rules

7.7 L4 — Validation System

Purpose

  • Ensure structural integrity

Checks

- contradictions
- missing definitions
- circular dependencies
- incomplete logic

Output

validation_report = {
  errors,
  warnings,
  unresolved_elements
}

Failure condition

  • unresolved critical inconsistencies

7.8 L5 — Failure System

Purpose

  • Actively identify breakpoints

Tests

- ambiguity
- contradiction
- omission
- Goodhart / optimization
- drift
- governance gaps

Output

failure_report = {
  type,
  location,
  impact,
  suggested_fix
}

Enforcement

  • system must attempt to break itself

Failure condition

  • absence of detectable failure in non-trivial input

7.9 L6 — Audit System

Purpose

  • Record full transformation trace

Structure

trace = [
  {step, input, output, timestamp, reason}
]

Enforcement

  • complete logging
  • replayability

Failure condition

  • missing or partial trace

7.10 L7 — Attribution System

Purpose

  • Assign responsibility for outputs

Output

decision_path = {
  source_claims,
  transformations,
  system_actions,
  human_overrides
}

Enforcement

  • no anonymous outputs

Failure condition

  • unassigned responsibility

7.11 L8 — Governance System

Purpose

  • Define control and escalation

Components

- permission model
- override rules
- escalation hierarchy
- human review checkpoints

Enforcement

  • governance must be explicit and enforced

Failure condition

  • implicit or bypassable authority

7.12 L9 — Optimization Boundary

Purpose

  • constrain optimization behavior

Functions

- detect metric gaming
- enforce bounded optimization
- expose tradeoffs

Enforcement

  • optimization cannot override constraints

Failure condition

  • unbounded optimization

7.13 L10 — Safety Core

Purpose

  • final gate before output

Checks

- validation complete
- no critical conflicts
- required fields present

Enforcement

  • block output if conditions fail

Failure condition

  • unsafe or incomplete outputs released

7.14 Data Flow Summary

Artifact
  ↓
Claims
  ↓
Provenance
  ↓
Constraints
  ↓
Validation
  ↓
Failure Analysis
  ↓
Audit + Attribution
  ↓
Governance
  ↓
Safety Gate
  ↓
Output

7.15 Architectural Properties


Deterministic Core

  • structure, constraints, and validation are deterministic where possible

Probabilistic Assistance

  • LLMs assist in:
    • claim extraction
    • ambiguity identification
  • but do not define truth

Composability

  • layers can be extended without breaking core invariants

Domain-Agnostic Design

  • same architecture applies across:
    • law
    • contracts
    • policy
    • investigative data
    • AI governance

7.16 Failure Modes of the Architecture

If misapplied:

- skipping layers → hidden ambiguity
- weak claim extraction → invalid constraints
- missing provenance → unverifiable outputs
- absent failure system → false confidence

7.17 Architectural Compression

Substrate =

structured input
→ explicit claims
→ enforced constraints
→ validated system
→ adversarial testing
→ auditable output

This architecture defines the operational core of Substrate.


8. Failure Harness

The Failure Harness is the defining component of Substrate. It converts the system from a passive analyzer into an active adversary of its own outputs.


8.1 Philosophy

Traditional systems ask:

Is this correct?

Substrate asks:

Where does this break?

The harness assumes:

  • inputs are incomplete
  • actors are strategic
  • definitions are unstable
  • systems will be exploited

8.2 Design Principle

Failure must be induced, not waited for.

The system must:

  • generate adversarial conditions
  • probe edge cases
  • simulate misuse
  • expose hidden assumptions

8.3 Failure Categories


8.3.1 Ambiguity

Detects undefined or non-operational language.

Examples:
- "reasonable"
- "appropriate"
- "as necessary"

Output:

Ambiguity detected → requires parameterization or definition

8.3.2 Contradiction

Identifies conflicting constraints.

Rule A: must provide X  
Rule B: may deny X under condition Y  

Output:

Conflict graph + resolution requirement

8.3.3 Omission

Detects missing elements.

Missing:
- enforcement mechanism
- authority
- definition
- scope boundary

Output:

Incomplete system warning

8.3.4 Goodhart / Optimization

Simulates incentive-driven exploitation.

Metric: maximize compliance  
Exploit: redefine compliance narrowly  

Output:

Metric divergence detected

8.3.5 Drift

Tests temporal stability.

- interpretation changes
- thresholds become outdated

Output:

Drift risk + recalibration requirement

8.3.6 Governance Failure

Tests authority and decision clarity.

- multiple authorities
- no escalation path

Output:

Authority conflict detected

8.4 Test Structure

Each test follows a standard format:

Test ID:
Failure Type:
Input:
Attack / Scenario:
Expected Failure:
System Requirement:
Pass Condition:
Fail Condition:

8.5 Example (Condensed)

Test: F2.1 — Contradiction

Input:
- must provide service
- may deny service if condition X

Attack:
- condition X is broadly interpreted

Expected Failure:
- contradictory enforcement

Pass:
- system flags conflict

Fail:
- system silently chooses one path

8.6 Refusal as Valid Output

A critical feature:

The system must be able to refuse to act.

Conditions for refusal:

  • unresolved ambiguity
  • missing definitions
  • conflicting constraints
  • incomplete validation

This prevents:

false precision → incorrect execution

8.7 Output of the Harness

failure_report = {
  failures: [
    {type, location, severity, explanation}
  ],
  risk_profile,
  suggested_fixes
}

8.8 Severity Levels

Critical:
- blocks execution

Major:
- produces inconsistent outcomes

Minor:
- reduces clarity or robustness

8.9 Integration with System

The harness operates after validation but before final output:

constraints → validation → failure harness → output gate

It can:

  • block outputs
  • trigger rewrite
  • escalate to governance layer

8.10 Adversarial Orientation

The harness assumes:

users, institutions, and systems will attempt to:
- exploit ambiguity
- minimize obligations
- maximize advantage

Thus it models:

  • worst-case behavior
  • not ideal usage

8.11 Harness Properties


Systematic

Covers all major failure classes.


Repeatable

Tests can be rerun across versions.


Extensible

New failure modes can be added.


Domain-agnostic

Applies across:

  • law
  • contracts
  • policy
  • AI systems

8.12 Compression

Failure Harness =

structured adversarial testing
→ identifies where system breaks
→ forces explicit fixes or refusal

The harness transforms Substrate from a descriptive system into a defensive and diagnostic system.


9. Applied Example (Legislation)

This section demonstrates Substrate applied to real legislative patterns using two representative cases:

  1. ADA-style statutory language (labor / accommodation)
  2. Healthcare coverage policy (ACA-style provision)

The objective is not legal interpretation, but structural transformation:

original text → claims → constraints → failure detection → rewrite (v0.2)

9.1 Example A — ADA-Style Provision

Original Form (Condensed)

Employers must provide reasonable accommodations to qualified individuals unless doing so would impose undue hardship.

9.2 Claim Extraction

C1: employer must provide accommodation
C2: applies to qualified individuals
C3: exception exists (undue hardship)

9.3 Constraint Construction

IF individual is qualified
AND accommodation is required
THEN employer must provide accommodation
UNLESS undue hardship

9.4 Failure Detection

Ambiguity

- "reasonable"
- "qualified"
- "undue hardship"

→ non-operational terms


Constraint Gap

No definition of:
- qualification criteria
- hardship threshold

Incentive Risk

Employer incentive:
- redefine "qualified"
- overstate hardship

Governance Gap

No explicit authority:
- who determines hardship?

9.5 Failure Summary

- undefined key terms
- missing thresholds
- incentive misalignment
- unclear authority

9.6 Rewritten Version (v0.2)

ELIGIBILITY:
Individual qualifies if:
- meets documented job requirements at hiring
- OR meets justified, externally reviewable updates

ACCOMMODATION:
Defined as:
- physical modification
- schedule adjustment
- assistive technology
- OR functionally equivalent intervention

UNDUE HARDSHIP:
TRUE if:
- cost > X% revenue (standardized)
- OR operational disruption > threshold Y

REQUIREMENTS:
- all hardship claims must include evidence
- all decisions must be logged and auditable

GOVERNANCE:
- disputes escalate: employer → regulator → court

9.7 Transformation Summary

ElementBeforeAfter
Termsambiguousdefined/parameterized
Constraintsimplicitexplicit
Exceptionsvaguebounded
Governanceimplicitstructured
Incentiveshiddenexposed

9.8 Example B — Healthcare Coverage (ACA-Style)

Original Form

A health insurance issuer shall provide coverage for preventive services. Services may be denied if not medically necessary.

9.9 Claim Extraction

C1: issuer must provide preventive services
C2: services defined externally
C3: denial allowed if not medically necessary

9.10 Constraint Construction

IF service ∈ preventive
THEN coverage required

EXCEPTION:
IF not medically necessary
THEN denial allowed

9.11 Failure Detection

Contradiction

preventive ≠ medically necessary (always)

→ rule conflict


Ambiguity

- "preventive"
- "medically necessary"

Authority Conflict

- external body defines services
- agency may override

Incentive Risk

issuer minimizes cost by:
- redefining necessity

9.12 Failure Summary

- conflicting constraints
- undefined definitions
- dual authority
- denial bias

9.13 Rewritten Version (v0.2)

PREVENTIVE SERVICES:
- defined by authority A
- OR approved by Secretary under criteria

MEDICAL NECESSITY:
- measurable health outcome improvement
- OR defined probability threshold

CONFLICT RULE:
- latest valid definition applies
- all changes versioned

DENIAL CONDITIONS:
- must include evidence
- must be logged and reviewable

AUDIT:
- abnormal denial rates trigger review

APPEALS:
- independent review required

9.14 Key Observations


Single Sentence → Multi-Claim System

1 sentence → 3–5 claims → multiple constraints

Failure Is Immediate

Even minimal text produces:

  • ambiguity
  • conflict
  • omission

Rewrite Is Structural, Not Stylistic

Changes are:

  • definitional
  • parametric
  • governance-oriented

9.15 General Pattern

Across both examples:

text → decomposition → constraint mapping → failure detection → structured rewrite

9.16 Result

Substrate produces:

- failure report
- structured system
- explicit uncertainty
- improved operational form

9.17 Insight

These examples demonstrate:

Laws are typically policy-complete but system-incomplete

Substrate makes them:

system-complete but explicitly constrained

This section establishes the system’s practical operation on real legislative structures.


10. What Changed (Critical Insight)

This section isolates the transformation Substrate performs. It is not a marginal improvement in drafting clarity; it is a change in the type of artifact being produced.


10.1 From Language to System

Before Substrate:

law = language artifact

After Substrate:

law = structured, constrained, testable system

10.2 Transformation Axes


10.2.1 Interpretation → Explicit Structure

Before:

  • meaning distributed across:
    • text
    • courts
    • agencies
    • practice

After:

- claims extracted
- dependencies mapped
- interpretation points surfaced

Result:

interpretation becomes a visible layer, not a hidden process

10.2.2 Ambiguity → Surfaced and Localized

Before:

  • ambiguity embedded in text
  • often discovered through litigation

After:

ambiguity_report = {
  term,
  location,
  impact,
  required_definition
}

Result:

ambiguity is detected at draft-time, not post-enactment

10.2.3 Constraints → Explicit and Testable

Before:

  • rules expressed rhetorically

After:

IF condition
THEN obligation
UNLESS exception

Result:

rules become evaluable structures

10.2.4 Exceptions → Bounded

Before:

  • broad exceptions (“unless undue hardship”)

After:

exception = {
  threshold,
  measurement,
  evidence_required
}

Result:

exceptions become constraints, not escape hatches

10.2.5 Governance → Structured

Before:

  • authority implicit
  • escalation undefined

After:

authority_chain = [
  internal_actor,
  regulator,
  judiciary
]

Result:

decision pathways become explicit

10.2.6 Incentives → Visible

Before:

  • incentive effects implicit
  • discovered through behavior

After:

risk_profile = {
  incentive_vector,
  exploit_path,
  mitigation
}

Result:

system anticipates adversarial use

10.2.7 Drift → Trackable

Before:

  • changes accumulate across:
    • amendments
    • interpretation
    • enforcement

After:

versioned_rule = {
  t0,
  t1,
  t2,
  deltas
}

Result:

temporal evolution becomes analyzable

10.2.8 Validation → Systematic

Before:

  • correctness evaluated informally

After:

validation = {
  contradictions,
  omissions,
  dependencies,
  completeness
}

Result:

integrity becomes measurable

10.2.9 Failure → First-Class Output

Before:

  • failure discovered through:
    • litigation
    • enforcement breakdown

After:

failure_report = primary output

Result:

system identifies failure before deployment

10.3 Compression

Before:
- language
- interpretation
- enforcement
- litigation

After:
- structure
- constraints
- validation
- failure detection
- audit

10.4 Key Shift

The central transformation is:

policy intent → operational system

Not:

  • clearer wording
  • better summaries
  • improved drafting style

10.5 What Is Not Changed

Substrate does not:

  • remove ambiguity entirely
  • eliminate interpretation
  • determine correct meaning
  • resolve disputes
  • replace institutions

10.6 What Is Changed

Substrate changes:

- when ambiguity is detected (earlier)
- where it is located (explicitly)
- how it is handled (structured)
- who can see it (broader access)

10.7 Outcome

The system shifts from:

reactive (litigation-driven)

to:

proactive (structure-driven)

10.8 Final Statement

Substrate does not make law correct.

It makes law inspectable, testable, and accountable to its own structure.

11. Domain Generalization

Substrate is not specific to law. It applies to a broader class of systems defined by:

language → interpreted as rules → producing real-world consequences

This section organizes applicable domains into archetypes, each requiring minimal adaptation of the core architecture.


11.1 Archetype Overview

1. Normative Systems
2. Institutional Decision Systems
3. Economic / Incentive Systems
4. Investigative / Evidence Systems
5. AI Governance Systems
6. Scientific / Knowledge Systems

Each archetype shares:

  • structured or semi-structured language
  • interpretation layers
  • non-trivial consequences
  • hidden failure modes

11.2 Normative Systems (Rules / Governance)

Examples

  • legislation
  • regulation
  • contracts
  • standards (ISO, IEEE)
  • corporate policies

Fit

claims → constraints → validation → failure

Substrate role

  • compile rules into constraints
  • detect contradictions and omissions
  • expose governance gaps

Primary outputs

  • conflict maps
  • ambiguity reports
  • enforceability analysis

11.3 Institutional Decision Systems

Examples

  • courts
  • administrative agencies
  • corporate governance
  • university policies

Hidden structure

written rules + implicit decision criteria

Substrate role

  • map decision rules to explicit claims
  • detect inconsistency across decisions
  • surface implicit criteria

Primary outputs

  • decision consistency analysis
  • interpretation divergence
  • drift over time

11.4 Economic / Incentive Systems

Examples

  • tax systems
  • subsidies
  • tariffs, trade rules
  • procurement
  • benefit programs

Core issue

rules + incentives → behavior

Substrate role

  • simulate Goodhart effects
  • detect exploit paths
  • map beneficiary structures (structurally, not normatively)

Primary outputs

  • incentive risk profiles
  • exploit scenarios
  • structural benefit flows

11.5 Investigative / Evidence Systems

Examples

  • Panama Papers
  • Epstein corpus
  • FOIA archives
  • financial disclosures

Shift

rules → relationships

Substrate role

  • extract claims from documents
  • build relationship graphs
  • detect contradictions and omissions
  • preserve provenance

Primary outputs

  • entity graphs
  • timeline inconsistencies
  • cross-document linkages

11.6 AI Governance Systems

Examples

  • model policies
  • safety rules
  • evaluation frameworks
  • alignment specifications

Problem

vague rules → inconsistent enforcement

Substrate role

  • formalize policy constraints
  • stress-test failure modes (jailbreaks, edge cases)
  • detect incentive misalignment

Primary outputs

  • policy constraint maps
  • failure scenarios
  • enforcement gaps

11.7 Scientific / Knowledge Systems

Examples

  • research papers
  • clinical guidelines
  • protocols
  • meta-analyses

Problem

claims without consistent validation or reproducibility

Substrate role

  • extract claims and dependencies
  • check internal consistency
  • identify unsupported assertions

Primary outputs

  • claim dependency graphs
  • contradiction detection across studies
  • missing evidence flags

11.8 Cross-Archetype Invariants

Across all domains, Substrate performs the same functions:

1. make implicit structure explicit
2. detect contradictions
3. expose incentives
4. surface uncertainty
5. enable adversarial queries

11.9 Domain Adaptation Layer

Each domain requires:

- domain-specific claim types
- domain-specific constraints
- domain-specific failure tests

Core architecture remains unchanged.


11.10 Non-Applicable Domains

Substrate is not effective for:

- purely formal systems (e.g., mathematics)
- purely creative systems (e.g., fiction)
- systems without interpretive layers

11.11 Compression

Substrate applies wherever:
rules + language + incentives + uncertainty intersect

11.12 Implication

Substrate is not a domain-specific tool.

It is a general system for exposing hidden structure in interpretation-dependent systems.


Substrate scales from a single drafting/validation system into a federated network of legal and policy nodes, each representing a jurisdiction or institutional system.


12.1 Core Principle

Do not centralize law.
Expose structure across jurisdictions.

Substrate does not merge legal systems into a single authority. It connects them while preserving sovereignty.


12.2 Node Definition

Each node represents a legal system:

Node = {
  jurisdiction,
  legal corpus,
  claims,
  constraints,
  interpretations,
  provenance
}

Examples:

  • U.S. federal law
  • state law
  • EU directives
  • international treaties
  • regulatory regimes

12.3 Node Interfaces

Nodes expose:

- structured claims
- constraint graphs
- definitions
- version history
- interpretation overlays

This allows cross-node comparison without altering internal authority.


12.4 Cross-Node Relationships

Substrate enables:

- equivalence (same concept, different wording)
- conflict (incompatible constraints)
- dependency (one system references another)
- override (hierarchical authority)

12.5 Example: Trade Systems

Country A: tariff classification X  
Country B: tariff classification Y  

Result:
- inconsistent import rules
- compliance conflict

Substrate output:

Conflict map:
- classification mismatch
- operational consequence

12.6 Sovereignty Preservation

Each node retains:
- authority
- interpretation
- enforcement

Substrate does not:

  • resolve conflicts
  • impose harmonization

It reveals:

where harmonization is required or impossible

12.7 Federation Model

Conceptually similar to:

X-Road (data exchange layer)
+
Substrate (constraint/meaning layer)

Transport ≠ meaning.

Substrate operates at the semantic and constraint level.


12.8 Benefits


Cross-Jurisdiction Visibility

- conflicting obligations
- redundant rules
- incompatible standards

Policy Alignment

- identify harmonization opportunities
- detect structural divergence

Risk Detection

- multi-jurisdiction compliance gaps
- trade and regulatory conflicts

12.9 Scaling Strategy

Federation grows by:

node → node → node → network

Not by centralizing all data at once.


12.10 Minimal Version

single jurisdiction
→ structured claims
→ local validation

12.11 Expanded Version

multiple jurisdictions
→ cross-node comparison
→ conflict detection

12.12 Full Network

global nodes
→ linked constraints
→ interpretation overlays
→ drift tracking

12.13 Key Property

Substrate does not unify law.
It makes differences explicit.

12.14 Compression

Federated Legal Constraint Network =

independent nodes
+ shared structure
+ exposed differences
+ preserved authority

13. Interpretation Layer

The Interpretation Layer models how rules are actually applied across jurisdictions and over time. It does not decide correctness; it records, structures, and compares interpretations.


13.1 Role

text → interpretation → application

Substrate represents interpretation as a first-class overlay on claims and constraints.


13.2 Interpretation Object

interpretation = {
  claim_id,
  authority,          // court, agency, body
  jurisdiction,
  date,
  scope,              // broad / narrow / neutral
  effect,             // expands, restricts, distinguishes
  holding_strength,   // binding, persuasive
  notes,
  provenance
}

13.3 Attachment to Claims

claim C1
  ├─ interpretation A (9th Circuit, expands scope)
  ├─ interpretation B (5th Circuit, restricts scope)
  └─ interpretation C (agency guidance, narrows conditions)

13.4 Delta Detection (Conflict Mapping)

delta = {
  claim_id,
  interpretations: [A, B],
  difference_type: scope_conflict,
  impact: compliance divergence
}

Outputs:

  • conflict graphs
  • jurisdictional divergence maps

13.5 Circuit / Jurisdiction Splits

jurisdiction_map = {
  9th Circuit: X
  5th Circuit: Y
  2nd Circuit: Z
}

Result:

  • explicit, queryable “split” representation
  • no resolution implied

13.6 Temporal Drift

timeline = [
  {date: t0, interpretation: A},
  {date: t1, interpretation: B},
  {date: t2, interpretation: C}
]

Detects:

  • silent narrowing/expansion
  • instability zones

13.7 Uncertainty Mapping

uncertainty = {
  claim_id,
  interpretation_count,
  conflict_density,
  stability_score
}

High conflict density ⇒ low certainty.


13.8 Authority Hierarchy (Non-Resolving)

hierarchy = [
  trial_court,
  appellate_court,
  supreme_court,
  agency
]

Used to:

  • contextualize interpretations
  • not to enforce a single “correct” outcome within Substrate

13.9 Inputs (Phased)

  • Phase 1: manual / curated mappings
  • Phase 2: structured extraction from opinions
  • Phase 3: automated clustering and linkage

13.10 Outputs

- interpretation overlays
- conflict graphs
- drift timelines
- uncertainty maps

13.11 Constraints

  • no case outcome prediction
  • no legal advice
  • no ranking of correctness beyond structural signals (e.g., binding vs persuasive)

13.12 Compression

Interpretation Layer =

claims
+ jurisdictional overlays
+ temporal evolution
+ conflict detection
→ structured uncertainty

14. System Boundaries

Substrate’s value depends on clear, enforced boundaries. These are not disclaimers alone; they are design constraints that shape system behavior, outputs, and governance.


14.1 Non-Functions (Explicitly Excluded)

Substrate does not:

- provide legal advice
- predict case outcomes
- recommend litigation strategy
- determine legal rights or liabilities
- replace licensed professionals (lawyers, judges)
- assert normative conclusions (e.g., “this is fair/unfair”)
- collapse conflicting interpretations into a single truth

These exclusions are enforced at:

  • output generation (no prescriptive claims)
  • UI (no “you should…” language)
  • governance (review and escalation)

14.2 Positive Functions (What It Does)

Substrate does:

- extract and structure claims from language
- build constraint and dependency graphs
- detect ambiguity, contradictions, omissions
- map interpretations across jurisdictions and time
- expose incentive risks and exploit paths (structurally)
- produce auditable traces and provenance links
- generate uncertainty maps and failure reports

14.3 Output Constraints

All outputs must satisfy:

- source-linked (provenance attached)
- uncertainty-explicit (no hidden assumptions)
- non-prescriptive (no advice or directives)
- auditable (traceable transformations)
- reversible (replayable steps)

Prohibited output patterns:

- “This will win in court”
- “You should do X to comply”
- “This is legally valid/invalid” (without framing as structural finding)

Allowed framing:

- “Undefined term: ‘reasonable’ (location, impact)”
- “Conflict between Rule A and Rule B (graph attached)”
- “High uncertainty due to divergent interpretations (map attached)”

14.4 User Interaction Boundaries

Substrate supports:

- drafting assistance (structure + failure detection)
- comparison across laws/policies
- exploratory queries (where does this conflict? what is undefined?)

Substrate does not support:

- personalized legal guidance
- jurisdiction-specific advice tailored to a user’s facts
- decision recommendations with legal consequences

When user intent approaches advice, the system must:

- reframe to structural analysis
- surface uncertainty
- suggest consulting qualified professionals

14.5 Interpretation Boundaries

Substrate records interpretations but does not resolve them:

- shows circuit/jurisdiction splits
- indicates binding vs persuasive authority
- displays temporal drift

It does not:

- declare a “correct” interpretation
- rank courts beyond structural signals
- predict how a specific case will be decided

14.6 Incentive and Outcome Boundaries

Substrate may:

- expose structural relationships (e.g., rule → potential beneficiary class)
- link to public data (budgets, contracts) when available

It must not:

- assert causation (e.g., “X caused Y”)
- allege wrongdoing
- draw defamatory conclusions

Outputs are framed as:

- “linked entities and events”
- “observable patterns”
- “open questions for investigation”

14.7 Data and Provenance Constraints

All claims and links require provenance:

claim → source → section → version → timestamp

The system must:

- block or flag claims without sources
- distinguish primary vs derived data
- preserve document integrity (no silent alteration)

14.8 Safety Core Enforcement

Before release, outputs pass a final gate:

block if:
- critical validation failures unresolved
- provenance missing
- outputs drift into advice/prediction
- governance/attribution incomplete

14.9 Governance and Accountability

- actions attributable (system + human)
- overrides logged
- escalation paths defined

No anonymous decisions:

every output → decision_path

The boundary is implemented through behavior:

- analytical, not advisory
- descriptive, not prescriptive
- structural, not determinative

This alignment supports a defensible posture:

Substrate provides educational, analytical, and structural insights with explicit uncertainty and provenance, and does not substitute for professional judgment.

14.11 Compression

Substrate does not decide.
It reveals structure, conflict, and uncertainty.

15. Institutional Reaction

Substrate introduces structural transparency into systems that are traditionally opaque, interpretive, and distributed. Institutional reaction is therefore mixed, phased, and incentive-dependent.


15.1 Reaction Phases


Phase 1 — Acceptance (Low Threat)

Use cases:

  • citizen drafting assistance
  • ambiguity detection
  • conflict identification
  • structural comparison

Perception:

“Helpful tool for clarity and participation”

Likely supporters:

  • civic tech organizations
  • academic institutions
  • policy researchers
  • journalists
  • standards bodies

Phase 2 — Selective Adoption

Use cases:

  • internal policy drafting
  • compliance structuring
  • regulatory analysis

Perception:

“Useful internally, but sensitive externally”

Behavior:

  • adopted behind institutional boundaries
  • limited public exposure
  • partial integration into workflows

Phase 3 — Resistance

Triggered when system exposes:

- legislative contradictions
- enforcement gaps
- incentive misalignment
- structural beneficiaries

Perception:

“Destabilizing / politically sensitive”

Likely resistance from:

  • political actors
  • regulatory bodies
  • organizations benefiting from opacity

Phase 4 — Normalization

Over time:

- structural transparency becomes expected
- drafting standards evolve
- institutions adapt to visibility

Comparable shift:

  • financial disclosure systems
  • open data initiatives

15.2 Stakeholder Analysis


Legislatures

Positive

  • improved drafting quality
  • reduced unintended consequences

Concerns

  • exposure of internal trade-offs
  • reduced flexibility in language

Courts

Positive

  • structured view of statutory ambiguity
  • visibility of interpretive divergence

Concerns

  • perceived encroachment on interpretive authority

Constraint:

system must not decide cases or predict outcomes

Regulatory Agencies

Positive

  • clearer enforcement frameworks
  • audit and compliance support

Concerns

  • exposure of enforcement inconsistency
  • increased scrutiny

Corporations / Regulated Entities

Positive

  • clearer compliance mapping
  • reduced ambiguity risk

Concerns

  • reduced ability to exploit ambiguity
  • increased accountability

Journalists / Researchers

Positive

  • structured investigative tool
  • traceable relationships and gaps

Concerns

  • data interpretation responsibility

Public / Citizens

Positive

  • increased access to structure
  • improved participation in drafting

Concerns

  • complexity of outputs

15.3 Incentive Alignment

Institutional reaction depends on:

benefit_from_clarity vs benefit_from ambiguity

Actors benefiting from:

  • clarity → adopt
  • ambiguity → resist

15.4 Adoption Path

Substrate is most likely to succeed by:

1. starting with low-risk use cases
2. demonstrating utility in drafting and validation
3. expanding into analysis and comparison
4. gradually exposing deeper structure

15.5 Governance Implications

To sustain adoption:

- neutrality must be maintained
- outputs must remain non-prescriptive
- provenance must be explicit
- uncertainty must be visible

15.6 Risk Factors


Misuse

  • outputs misinterpreted as advice
  • structural findings framed as conclusions

Overreach

  • expansion into prediction or decision-making
  • erosion of boundary constraints

Data Sensitivity

  • linking datasets across domains
  • exposure of sensitive relationships

15.7 Mitigation

- strict system boundaries
- output framing controls
- audit and traceability
- phased rollout

15.8 Compression

Institutions support transparency in principle.
They resist it when it becomes operational.

15.9 Final Observation

Substrate does not remove institutional authority.

It changes:

what is visible
when it becomes visible
who can access it

16. Development Phases

Substrate must be built incrementally. The architecture supports large-scale integration, but the system only becomes viable through phased construction, where each phase produces standalone value.


16.1 Phase Design Principle

Each phase must:
- produce usable output
- not depend on future phases
- preserve compatibility with expansion

Avoid:

“build everything, then release”

16.2 Phase 1 — Law Compiler (Foundational)

Scope

  • ingest legal text
  • extract claims
  • build basic constraints
  • run validation + failure harness

Capabilities

draft → claims → constraint check → failure report → rewrite

Data Required

  • existing federal law (static corpus sufficient)

Outputs

  • ambiguity detection
  • contradiction detection
  • omission flags
  • structured rewrite

Value

  • immediate utility for drafting
  • no dependency on external systems

16.3 Phase 2 — Process Linking

Scope

  • connect bills, amendments, sponsors, votes

Capabilities

bill → versions → changes → sponsors → voting record

Data Sources

  • Congress APIs
  • GovInfo
  • GovTrack / ProPublica Congress

Outputs

  • version diffs
  • amendment impact mapping
  • sponsor linkage

Value

  • transparency of legislative evolution
  • early process insight

16.4 Phase 3 — Influence Overlay

Scope

  • connect funding, lobbying, and voting patterns

Capabilities

bill → sponsors → funding → votes → alignment patterns

Data Sources

  • FEC (campaign finance)
  • IRS 990 (nonprofits)
  • OpenSecrets (aggregated)
  • voting records

Outputs

  • structural relationships (not causal claims)
  • potential influence patterns
  • queryable linkages

Constraint

no assertion of causation

16.5 Phase 4 — Outcome Linking

Scope

  • connect law to resource allocation and execution

Capabilities

law → budget → contracts → distribution

Data Sources

  • USAspending.gov
  • OMB / budget data
  • agency datasets

Outputs

  • structural mapping of resource flow
  • linkage between rules and allocations

Constraint

no normative or causal claims

16.6 Phase 5 — Drift Engine

Scope

  • track changes over time

Capabilities

rule(t0) → rule(t1) → rule(t2)

Functions

  • version tracking
  • amendment diffing
  • interpretation drift (later phase)

Outputs

  • drift timelines
  • stability/instability zones

16.7 Phase 6 — Query System

Scope

  • enable structured, cross-layer queries

Capabilities

“show conflicts”
“show changes”
“show relationships”

Outputs

  • multi-layer query results
  • graph-based exploration

Example Queries

- Where does this bill conflict with existing law?
- How did this provision change over time?
- Which rules have the highest ambiguity?

16.8 Phase 7 — Interpretation Layer Integration

Scope

  • attach court and agency interpretations

Capabilities

claim → interpretations → conflict maps

Data Sources

  • CourtListener
  • judicial corpora

Outputs

  • circuit splits
  • interpretation overlays
  • uncertainty maps

16.9 Phase 8 — Federated Network

Scope

  • multi-jurisdiction integration

Capabilities

federal ↔ state ↔ international ↔ treaties

Outputs

  • cross-jurisdiction conflict maps
  • harmonization opportunities
  • compliance divergence

16.10 Phase Compression

Phase 1: structure
Phase 2: process
Phase 3: influence
Phase 4: outcomes
Phase 5: time
Phase 6: queries
Phase 7: interpretation
Phase 8: federation

16.11 Critical Constraint

Do not skip phases.

Each phase introduces a new dimension:

  • structure
  • relationships
  • incentives
  • execution
  • time
  • interpretation
  • federation

Skipping introduces instability.


16.12 Practical Recommendation

Focus initial build on:

Phase 1 + partial Phase 2

This provides:

  • immediate value
  • manageable complexity
  • strong foundation

16.13 Compression

Substrate evolves from:
draft validation system
→ legislative analysis system
→ governance mapping system
→ federated constraint network

17. Minimal Viable System

Substrate’s minimal viable system (MVS) is intentionally narrow. It does not attempt to implement the full architecture or all phases. It focuses on a single, high-value loop:

draft → structure → detect failure → return report → suggest rewrite

17.1 Scope

The MVS operates on:

- existing law (static corpus)
- user-submitted draft text

It does not require:

  • lobbying data
  • budget data
  • contracts
  • court interpretation layers
  • cross-jurisdiction federation

17.2 Core Components


17.2.1 Ingestion

input:
- draft proposal
- reference law corpus

Requirements:

  • preserve text
  • segment into clauses

17.2.2 Claim Extraction

text → atomic claims

Output:

claim = {
  subject,
  action,
  condition,
  exception
}

17.2.3 Constraint Builder

claims → IF/THEN/UNLESS structures

Output:

  • basic constraint graph

17.2.4 Validation Engine

Checks:

- undefined terms
- missing dependencies
- incomplete rules

17.2.5 Failure Harness (Reduced)

Includes:

- ambiguity detection
- contradiction detection
- omission detection

Excludes (for MVS):

  • Goodhart simulation
  • drift modeling
  • governance layering (full)

17.2.6 Output Generator

Produces:

- ambiguity report
- conflict report
- missing elements
- structured rewrite suggestions

17.3 Example Output (MVS)

FAILURE REPORT

Ambiguity:
- “reasonable” (undefined)
- “appropriate” (undefined)

Missing:
- enforcement authority
- threshold definitions

Conflict:
- Rule A vs Rule B (details)

Suggested Rewrite:
- define threshold X
- specify authority Y

17.4 UI / UX (Minimal)


Input

text box for draft
optional reference selection (law corpus)

Output Panels

1. Structured Claims
2. Failure Report
3. Suggested Rewrite

Interaction

  • click term → see ambiguity details
  • click conflict → see constraint graph
  • toggle original vs structured version

17.5 What MVS Excludes


No Interpretation Layer

- no court mapping
- no circuit splits

No Incentive Modeling

- no lobbying data
- no funding links

No Outcome Linking

- no budget or contract data

No Federation

- single jurisdiction only

17.6 Why This Is Sufficient

The MVS already provides:

- immediate drafting improvement
- early failure detection
- structural clarity
- reusable outputs

It addresses the core problem:

language treated as logic without structure

17.7 Performance Characteristics


Low Cost

  • small model usage
  • limited data footprint

High Signal

  • failure detection scales with complexity
  • even short text produces useful output

Fast Iteration

  • users can revise drafts quickly
  • feedback loop is immediate

17.8 Expansion Path

MVS naturally extends to:

+ version tracking (Phase 2)
+ interpretation overlays (Phase 7)
+ cross-jurisdiction comparison (Phase 8)

17.9 Risk Management

MVS avoids:

- legal advice classification
- overreach into prediction
- dependence on sensitive data

17.10 Compression

Minimal Substrate =

structure + validation + failure detection

applied to drafting

This system alone can operate independently and deliver sustained value while the broader architecture develops.


18. Infrastructure Overview

Substrate’s infrastructure must support structured transformation, traceability, and scalable analysis. The system is not primarily storage-bound; it is structure- and computation-bound, with cost driven by parsing, validation, and model-assisted steps.


18.1 Core Data Types

Substrate operates on a small set of fundamental objects:

- artifact (raw text)
- claim (atomic unit)
- constraint (logic structure)
- interpretation (overlay)
- trace (audit record)
- failure (diagnostic output)

These objects must be:

  • versioned
  • addressable
  • linkable

18.2 Storage Layers


18.2.1 Document Store

Purpose:

  • store raw legal/policy text

Examples:

  • object storage (S3-style)
  • document databases

Characteristics:

- append-only
- versioned
- immutable references

18.2.2 Structured Data Store

Purpose:

  • store claims, constraints, metadata

Examples:

  • relational DB (Postgres)
  • document DB (JSON-based)

Schema (simplified):

claim {
  id,
  subject,
  action,
  condition,
  exception,
  source_ref
}

18.2.3 Graph Layer

Purpose:

  • represent dependencies and relationships

Examples:

  • graph DB (Neo4j)
  • graph layer over relational DB

Used for:

- constraint graphs
- interpretation overlays
- cross-document linking

18.2.4 Vector / Retrieval Layer

Purpose:

  • assist with semantic lookup
  • retrieve relevant sections

Examples:

  • vector DB (Pinecone, pgvector)

Used for:

- locating relevant law sections
- clustering similar claims

18.3 Processing Layer


18.3.1 Parsing / Extraction

Functions:

- segmentation
- claim extraction
- type assignment

Implementation:

  • rule-based parsing
  • LLM-assisted extraction

18.3.2 Constraint Builder

Functions:

- convert claims → logic structures
- build dependency graph

18.3.3 Validation Engine

Functions:

- detect contradictions
- detect missing definitions
- check completeness

18.3.4 Failure Harness

Functions:

- run adversarial tests
- generate failure reports

18.4 Model Usage (LLM Layer)

LLMs are used as assistive components, not authoritative engines.


Roles

- claim extraction (from text)
- ambiguity identification
- rewrite suggestion generation

Constraints

- outputs must be validated structurally
- no direct trust in model outputs
- bounded use for cost control

Optimization Strategies

- use small models for extraction
- use larger models only for complex analysis
- cache repeated operations
- batch processing where possible

18.5 APIs and Data Sources (Phased)


Phase 1

- static law corpus (GovInfo bulk)

Phase 2

- Congress APIs
- GovTrack / ProPublica Congress

Phase 3

- FEC
- IRS 990
- OpenSecrets

Phase 4

- USAspending
- OMB / budget data

Phase 7

- CourtListener
- judicial corpora

18.6 Cost Drivers

Primary cost factors:


Model Usage

- number of analyses per user
- complexity of input (length, structure)
- model size used

Compute

- parsing and validation operations
- graph construction

Storage

- relatively low cost
- dominated by document size

- depends on query volume
- moderate cost

18.7 Cost Characteristics


Early Stage (MVS)

- low storage cost
- moderate model cost
- manageable compute

Scaled System

- model usage dominates
- graph queries increase
- cross-layer queries increase cost

18.8 Performance Strategy


Caching

- precompute law embeddings
- cache common queries

Incremental Processing

- process only changed sections
- reuse existing claims/constraints

Layer Separation

- decouple ingestion, processing, querying

18.9 Infrastructure Principles

- persistence over statelessness
- traceability over speed
- structure over raw text
- validation over generation

18.10 Compression

Substrate infrastructure =

document store
+ structured data layer
+ graph layer
+ retrieval layer
+ constrained model usage

This infrastructure supports both the minimal system and long-term expansion.


19. Limitations

Substrate is designed to expose structure, constraints, and failure. It does not eliminate the inherent complexity of language-based systems. This section defines the intrinsic limits of the approach.


19.1 Ambiguity Cannot Be Eliminated

Natural language contains irreducible ambiguity:

- contextual meaning
- evolving definitions
- domain-specific interpretation

Substrate can:

  • detect ambiguity
  • localize it
  • require definition

It cannot:

  • fully resolve ambiguity in all contexts

19.2 Interpretation Remains Necessary

Even with structured claims and constraints:

structure ≠ final meaning

Interpretation is required for:

  • edge cases
  • novel scenarios
  • evolving social and legal contexts

Substrate:

  • maps interpretations
  • exposes divergence

It does not:

  • replace interpretive authority

19.3 Incentives Cannot Be Fully Modeled

Substrate can identify structural incentive risks:

- Goodhart effects
- exploit paths

But cannot fully capture:

  • human behavior
  • political strategy
  • institutional dynamics

Thus:

incentive analysis = partial, not complete

19.4 Causality Is Not Determinable

When linking:

  • law
  • funding
  • outcomes

Substrate can show:

- correlations
- structural relationships
- temporal alignment

It cannot assert:

X caused Y

Causal inference remains outside system scope.


19.5 Data Quality Constraints

The system depends on:

- completeness of source data
- consistency across datasets
- accuracy of extraction

Risks:

  • missing data
  • inconsistent formats
  • outdated sources

Substrate must:

  • flag uncertainty
  • not assume completeness

19.6 Extraction Errors

Claim extraction relies partially on:

  • rule-based parsing
  • model-assisted interpretation

Errors may include:

- incorrect decomposition
- missed claims
- misclassified types

Mitigation:

  • validation layer
  • audit trace
  • iterative refinement

19.7 Scale Complexity

As system expands:

- graph size increases
- cross-layer interactions grow
- query complexity rises

Challenges:

  • performance
  • cost
  • interpretability of outputs

19.8 User Interpretation Risk

Users may:

- overinterpret outputs
- treat structural findings as conclusions
- misapply results

Mitigation:

  • explicit uncertainty
  • non-prescriptive outputs
  • clear boundary enforcement

19.9 Institutional Resistance

Limitations are not only technical:

- political sensitivity
- resistance to transparency
- selective adoption

System success depends on:

  • careful rollout
  • neutral positioning
  • adherence to boundaries

19.10 Non-Applicability Domains

Substrate is not suited for:

- purely formal systems (mathematics)
- purely creative systems (fiction)
- systems without rule interpretation layers

19.11 Partial System Visibility

Even with full implementation:

some system elements remain opaque:
- informal practices
- undocumented decisions
- implicit norms

Substrate reveals:

  • structured elements

Not:

  • all real-world behavior

19.12 Compression

Substrate improves visibility, not certainty.

It exposes:
- structure
- conflict
- uncertainty

It does not:
- resolve all ambiguity
- eliminate interpretation
- determine outcomes

19.13 Final Limitation Statement

Substrate cannot make complex systems simple.

It can make them legible.

20. Conclusion

Substrate is a system for transforming language-based institutional artifacts into structured, constrained, testable representations with explicit uncertainty and full provenance.

It does not attempt to replace:

  • legal systems
  • courts
  • regulators
  • experts

It does not attempt to:

  • decide cases
  • predict outcomes
  • eliminate ambiguity

Instead, it addresses a more fundamental gap:

institutional systems rely on language that is treated as logic,
without structure, validation, or failure detection

20.1 Core Contribution

Substrate introduces:

- claim decomposition
- constraint construction
- validation systems
- failure harness
- auditability
- interpretation mapping
- uncertainty surfacing

These elements convert:

language → inspectable system

20.2 Practical Value

At its minimal level, Substrate enables:

  • improved drafting
  • early failure detection
  • structural clarity

At scale, it enables:

  • cross-jurisdiction comparison
  • interpretation mapping
  • incentive exposure (structural)
  • temporal drift tracking
  • federated legal networks

20.3 System Identity

Substrate is best understood as:

a constraint-oriented analytical system

Not:

  • a legal advisor
  • a decision engine
  • a predictive model

20.4 Strategic Path

The system must be built:

incrementally

Starting with:

Phase 1: structure + validation + failure detection

Expanding to:

process → incentives → outcomes → interpretation → federation

20.5 Key Principle

do not collapse complexity
expose it

20.6 Final Statement

Substrate does not make institutional systems correct.

It makes their structure, conflicts, and uncertainties visible,
testable, and accountable.

APPENDICES


Appendix A — Canon (Full)

A system is reliable only if:
- claims are explicit
- constraints are enforceable
- validation is systematic
- outputs are auditable
- decisions are attributable
- failures are visible
- uncertainty is preserved

Appendix B — Primitive Set (Representative)

P1: claim decomposition
P2: subject/action separation
P3: condition extraction
P4: exception isolation
P5: constraint formation
P6: dependency mapping
P7: ambiguity detection
P8: contradiction detection
P9: omission detection
P10: provenance linking
P11: validation rules
P12: failure harness execution
P13: audit trace generation
P14: attribution assignment
P15: governance enforcement

Appendix C — Failure Test Templates

Test ID:
Failure Type:
Input:
Attack Scenario:
Expected Failure:
System Requirement:
Pass/Fail Conditions:

Appendix D — Example Structured Law (v0.2)

rule:
  eligibility: defined
  obligation: explicit
  exception: bounded
  governance: defined
  audit: required

Appendix E — Phase Data Sources

Phase 1:
- GovInfo (law corpus)

Phase 2:
- Congress APIs

Phase 3:
- FEC
- IRS 990
- OpenSecrets

Phase 4:
- USAspending
- OMB data

Phase 7:
- CourtListener

Appendix F — Minimal Data Model

claim {
  id,
  subject,
  action,
  condition,
  exception,
  source
}

constraint {
  id,
  logic,
  dependencies
}

trace {
  step,
  input,
  output,
  timestamp,
  reason
}

Canon v0.4 — Literature-Grounded System Canon

CANON-001 — Atomic Claim Requirement

A system cannot validate outputs unless claims are decomposed into explicit, inspectable units.
Status: stable
Source signal: legal/RAG papers repeatedly lacked atomic claims; CaseFacts partially confirmed claim-level validation.

CANON-002 — Validation ≠ Evaluation

Benchmarks, metrics, and LLM judges do not establish truth or correctness.
Status: stable
Source signal: legal RAG benchmarking, judicial extraction metrics, auditability literature.

CANON-003 — Constraint Must Be Enforced

A prompt, guideline, benchmark, or policy is not a constraint unless the system can prevent or detect violation.
Status: stable
Source signal: legal RAG, governance-constrained AI, assured autonomy.

CANON-004 — Provenance Is Necessary but Insufficient

Sources, citations, and logs are required, but do not by themselves prove correctness.
Status: stable
Source signal: legal datasets, RAG systems, audit frameworks.

CANON-005 — Formalization Enables Validation

Natural language alone cannot be reliably executed or checked; validation requires a structured or formal representation.
Status: stable
Source signal: computational law, Prolog/legal inconsistency, rules-as-code cluster.

CANON-006 — Translation Layer Is a Failure Surface

Every conversion from human language → machine representation → action introduces distortion risk.
Status: stable
Source signal: computational law, ontology distortion, assured autonomy.

CANON-007 — Ground Truth Can Fail

Validators, expert labels, benchmarks, and protected-group labels can be incomplete, noisy, biased, or outdated.
Status: stable
Source signal: legal RAG benchmark, noisy protected groups, fairness DRO.

CANON-008 — Validation Must Be Recursive

The validator must itself be validated.
Status: stable
Source signal: benchmark failure, auditability, recurring local validation.

CANON-009 — Local Validation Beats Global Assurance

One-time external validation cannot guarantee safety across sites, time, or deployment contexts.
Status: stable
Source signal: recurring local validation.

CANON-010 — Systems Drift

All deployed systems degrade under changing distributions, incentives, data pipelines, users, and environments.
Status: stable
Source signal: self-healing ML, recurring local validation, quantile activation.

CANON-011 — Diagnosis Must Precede Adaptation

Corrective action without root-cause diagnosis can worsen the system.
Status: stable
Source signal: self-healing ML.

CANON-012 — Optimization Creates Failure

Strong optimization selects for proxy error, tail risk, and gaming behavior.
Status: stable
Source signal: Goodhart papers, OSA paper.

CANON-013 — Objectives Are Not Fully Realizable

Specified objectives cannot perfectly encode intent, and trained systems do not perfectly satisfy specified objectives.
Status: stable
Source signal: Objective Satisfaction Assumption failure.

CANON-014 — Bounded Optimization Required

Systems must limit optimization pressure when the proxy-goal gap cannot be characterized.
Status: stable
Source signal: Goodhart / OSA.

CANON-015 — Worst-Case Reasoning Required

Safety cannot rely on average-case performance; systems must be tested against adversarial, shifted, and tail conditions.
Status: stable
Source signal: robust optimization, assured autonomy.

CANON-016 — Structure Beats Semantics

Semantic plausibility is not enough; operational systems require structural constraints: capacity, hierarchy, causality, temporality, state, and feasibility.
Status: stable
Source signal: assured autonomy, legal ontology/RAG, routing papers.

CANON-017 — Execution Traceability Required

Every material decision must be reconstructable from actual internal traces, not post-hoc explanation.
Status: stable
Source signal: Topaz, auditability, cloud evidence.

CANON-018 — Full-Fidelity Audit Required

Audit systems must capture complete, high-integrity execution evidence, not partial or lossy logs.
Status: stable
Source signal: saBPF, Advocate, auditability framework.

CANON-019 — Audit Evidence Must Be Accessible

Evidence must exist and be technically accessible to auditors through APIs, logs, monitoring, raw sources, or explainability tools.
Status: stable
Source signal: auditability framework.

CANON-020 — Audit Evidence Must Be Tamper-Resistant

Critical traces must be protected from retroactive alteration.
Status: stable
Source signal: blockchain audit, Advocate, zk audit.

CANON-021 — Privacy-Preserving Audit Is Possible

Auditability and confidentiality need not be opposites; zero-knowledge or aggregated evidence can verify without full disclosure.
Status: stable
Source signal: zk-MCP, Advocate.

CANON-022 — Responsibility Attribution Required

Multi-agent systems require attribution of which actor, model, component, or decision path caused an outcome.
Status: stable
Source signal: responsibility-gap audit system.

CANON-023 — Incentives Are Constraints

Systems are governed by incentives and power structures as much as formal rules.
Status: stable
Source signal: lobbying/delegation, peer review mechanism design.

CANON-024 — Adversarial Actors Are First-Class

Governance systems must assume strategic actors, gaming, omission, manipulation, and capture.
Status: stable
Source signal: KWBench, peer review, lobbying, Goodhart.

CANON-025 — Missing Information Is Signal

Omissions, absent stakeholders, missing provenance, and unavailable evidence must be treated as meaningful system signals.
Status: stable
Source signal: KWBench.

CANON-026 — Human Oversight Must Be Structured

Human-in-the-loop is not sufficient unless approval, override, escalation, and accountability are structurally enforced.
Status: stable
Source signal: governance-constrained agentic AI.

CANON-027 — Deterministic Safety Layer Required

Safety-critical components must be replayable, inspectable, and deterministic wherever possible.
Status: stable
Source signal: assured autonomy.

CANON-028 — Modularity Creates New Failure Surfaces

Agents, routers, RAG modules, validators, and orchestration layers each introduce independent failure modes.
Status: stable
Source signal: hybrid RAG, Topaz, agentic systems.

CANON-029 — Routing Must Be Explainable

If systems choose among models, tools, or agents, the selection logic must be visible and tunable.
Status: stable
Source signal: Topaz.

CANON-030 — Canon Must Remain Minimal

A canon should stabilize slowly, contain only reusable system rules, and avoid absorbing local observations as universal principles.
Status: enforced
Source signal: internal method / recovery pipeline.


Canon v0.4 Compression

The full canon reduces to one operational rule:

A system cannot be trusted unless claims are explicit, constraints are enforced, validation is recursive, evidence is auditable, optimization is bounded, failures are attributable, and governance is structurally embedded.

🧩 LAYER 0 — INPUT / ARTIFACT

Purpose

Capture raw material:

  • laws
  • claims
  • model outputs
  • policies
  • data

Canon mapping

  • CANON-001 (atomic claims)
  • CANON-004 (provenance required)

Requirement

  • No input enters system unstructured
  • Must be decomposable into units

🧩 LAYER 1 — CLAIM SYSTEM

Purpose

Convert artifacts into:

  • atomic claims
  • typed units
  • structured assertions

Canon mapping

  • CANON-001 (atomic claims)
  • CANON-005 (formalization)
  • CANON-016 (structure > semantics)

Enforcement

  • reject non-decomposable inputs
  • enforce claim schema

🧩 LAYER 2 — PROVENANCE SYSTEM

Purpose

Attach:

  • source
  • lineage
  • dependencies
  • temporal context

Canon mapping

  • CANON-004 (provenance necessary)
  • CANON-018 (full audit)
  • CANON-019 (accessible evidence)

Enforcement

  • no claim without provenance
  • provenance must be queryable + reconstructable

🧩 LAYER 3 — CONSTRAINT SYSTEM

Purpose

Define:

  • rules
  • invariants
  • boundaries
  • allowable states

Canon mapping

  • CANON-003 (enforced constraints)
  • CANON-015 (worst-case reasoning)
  • CANON-016 (structural constraints)

Enforcement

  • constraints must be:
    • executable
    • testable
    • non-bypassable

🧩 LAYER 4 — VALIDATION SYSTEM

Purpose

Test claims against:

  • constraints
  • evidence
  • other claims

Canon mapping

  • CANON-002 (validation ≠ evaluation)
  • CANON-008 (recursive validation)
  • CANON-009 (local validation)

Enforcement

  • must include:
    • negative testing
    • contradiction detection
    • multi-path validation

🧩 LAYER 5 — FAILURE SYSTEM

Purpose

Actively detect:

  • contradiction
  • drift
  • Goodhart effects
  • adversarial behavior

Canon mapping

  • CANON-010 (drift)
  • CANON-012 (optimization failure)
  • CANON-024 (adversarial actors)
  • CANON-025 (missing info as signal)

Enforcement

  • system must:
    • try to break itself
    • surface failure signals

🧩 LAYER 6 — AUDIT SYSTEM

Purpose

Capture:

  • full execution trace
  • decision paths
  • intermediate states

Canon mapping

  • CANON-017 (traceability)
  • CANON-018 (full audit)
  • CANON-020 (tamper resistance)

Enforcement

  • logs must be:
    • complete
    • immutable (or provably protected)
    • replayable

🧩 LAYER 7 — ATTRIBUTION SYSTEM

Purpose

Determine:

  • who/what caused an outcome
  • responsibility across components

Canon mapping

  • CANON-022 (responsibility attribution)
  • CANON-028 (modular failure surfaces)

Enforcement

  • every decision path must be attributable
  • no anonymous outcomes

🧩 LAYER 8 — GOVERNANCE SYSTEM

Purpose

Enforce:

  • permissions
  • human oversight
  • escalation paths
  • incentive alignment

Canon mapping

  • CANON-023 (incentives)
  • CANON-026 (structured human oversight)

Enforcement

  • governance must be:
    • structural
    • not optional
    • not bypassable

🧩 LAYER 9 — OPTIMIZATION SYSTEM (BOUNDARY)

Purpose

Control:

  • optimization pressure
  • tradeoffs
  • routing decisions

Canon mapping

  • CANON-013 (objective failure)
  • CANON-014 (bounded optimization)
  • CANON-029 (explainable routing)

Enforcement

  • must:
    • expose tradeoffs
    • prevent runaway optimization
    • operate under constraints

🧩 LAYER 10 — SAFETY CORE

Purpose

Final enforcement layer:

  • deterministic checks
  • system invariants
  • kill-switch conditions

Canon mapping

  • CANON-027 (deterministic safety layer)
  • CANON-015 (worst-case reasoning)

Enforcement

  • must be:
    • deterministic
    • minimal
    • independent of probabilistic components

🔁 SYSTEM FLOW (SIMPLIFIED)

Artifact
  ↓
Claim System
  ↓
Provenance
  ↓
Constraint System
  ↓
Validation
  ↓
Failure Detection
  ↓
Audit + Attribution
  ↓
Governance
  ↓
Optimization (bounded)
  ↓
Safety Core (final gate)

WHAT THIS ARCHITECTURE EXPLICITLY PREVENTS

  • silent hallucination
  • untraceable decisions
  • unverifiable outputs
  • unchecked optimization
  • hidden failure modes

IMPORTANT OBSERVATION

This is not how current systems are built.

Current systems:

input → model → output

Substrate:

input → structured system → constrained execution → validated output