Section 1 — Purpose, Scope, and Runtime Config Inventory

Purpose:

Convert repo reality into PR execution discipline.

For each BE / FE / scraper / infra structure:
- what is real now
- what is dormant
- what is mock / partial
- what is security-sensitive
- what is cost-sensitive
- what can be safely adapted now
- what must wait for a later stage

This document sits underneath the PR grids. The PR grids answer:

When should this structure be used?

This repo reality review answers:

What does this structure actually do today?
What settings can accidentally change behavior, cost, or validity?

Canonical rule:

A repo feature should not be treated as active system capability merely because code exists.

1. Runtime Config Inventory

This section identifies runtime settings that affect security, cost, validation integrity, hidden behavior, or stage-gating.

1.1 High-Signal Runtime Config Table

Config / SettingObserved Default / BehaviorUsed ForSecurity / Cost / Validity ConcernStage Guidance
SECRET_KEYDefaults to "your-secret-key-here" in backend settings.App security / signing.Dangerous if production/staging ever boots with default secret.Add startup validation before external users. Stage 0 can proceed if env is controlled.
JWT_SECRET_KEYDefaults to "your-secret-key".JWT signing.Same issue: default secret must never be accepted outside local/dev.Require non-default secret for Stage 6+.
ACCESS_TOKEN_EXPIRE_MINUTESDefaults to 60 * 24 * 8, about 8 days.Access token lifetime.Long token life increases risk if leaked.Acceptable for internal validation; revisit before external users.
SESSION_EXPIRY_DAYSDefaults to 7 days.Session lifetime.Same long-session risk.Review before Stage 6 external testing.
BACKEND_CORS_ORIGINSConfig default is ["*"], while main.py uses a hardcoded allowlist.Browser access control.Mixed source of truth. Future code may accidentally use wildcard config.Consolidate before external use. During Stage 0, verify active app uses allowlist.
CORS_METHODS, CORS_HEADERSConfig defaults are wildcard lists.CORS behavior.Less risky than wildcard origins, but still broad.Acceptable for dev; tighten once external users exist.
AALAM_MODELDefaults to "gpt-4".Aalam model selection.Cost-sensitive and validity-sensitive. Changing model between conditions invalidates comparisons.Every validation PR must record model/provider.
OPENAI_API_KEY, ANTHROPIC_API_KEY, GOOGLE_API_KEYOptional env-loaded provider keys.LLM provider access.Multiple providers can create hidden fallback/cost drift if code path is unclear.PRs must specify provider and confirm no fallback/parallel provider path.
RESERVOIR_LOAD_BOOT_DOCSDefaults to true.Reservoir boot-document preload.Major validation-contamination risk. Hidden context could influence outputs.Set false for PR9.5–PR11.5 unless explicitly testing reservoir.
ENABLE_RESERVOIR_PRELOADProperty mirrors RESERVOIR_LOAD_BOOT_DOCS.Same preload behavior.Same hidden-context risk.Treat as dormant until Stage 3+ or specific retrieval test.
RESERVOIR_BOOT_DOCS_BUSINESS_IDOptional env value.Scopes reservoir preload.If set, may change which documents load at boot.Must be recorded in any run manifest if reservoir is active.
RESERVOIR_URL, BOUNDARY_REGISTRY_URLDefault to localhost services.External/local service integration.May silently fail, timeout, or create startup side effects.Keep inactive for Stage 0 validation.
RESERVOIR_TIMEOUT, BOUNDARY_REGISTRY_TIMEOUTDefault 2 seconds.Service timeout.Low timeout may cause partial/degraded behavior; high timeout may slow tests.Record if testing integration/degradation.
QDRANT_URLDefaults to http://localhost:6334.Vector DB endpoint.If activated early, can introduce hidden retrieval.Dormant until PR15/19 observation or Stage 3 memory work.
QDRANT_TIMEOUTDefaults to 10 seconds.Vector DB timeout.Affects latency/cost of retrieval experiments.Only relevant once vector path is explicitly in scope.
REDIS_URL, MEMORY_BACKEND_URLDefaults to localhost Redis.Memory/cache/rate-limit potential.Redis availability can change runtime behavior.Record if used; do not assume in-memory and Redis behavior are equivalent.
RATE_LIMIT_PER_MINUTEConfig says 100/minute, while middleware has hardcoded endpoint limits.Rate/cost protection.Multiple rate-limit sources; middleware is in-memory and partially bypassed.For validation, use explicit run caps. For external use, use Redis-backed rate limits.
DEBUGDefaults false via env parse.Debug behavior.If true in production, may expose unsafe behavior/logging.Confirm false outside local dev.
POOL_SIZE, MAX_OVERFLOWDefaults 20 / 10.DB connection pool.Cost/resource-sensitive under multiple workers.Review before load/performance PR41.
WORKERSBackend Docker defaults to 4 workers.Uvicorn concurrency.Multiplies in-memory rate-limit capacity and concurrent LLM calls.For validation, prefer fixed known worker count; for scale, test explicitly.
NEXT_TELEMETRY_DISABLEDSet to 1 in frontend Dockerfile.Next.js telemetry.Good cost/privacy posture.Keep.
SKIP_ENV_VALIDATIONSet to 1 in frontend Docker build.Build-time env validation bypass.Can allow missing/misconfigured env to pass build.Revisit before product readiness.
NODE_ENV=productionSet in frontend Dockerfile.Production runtime mode.Good.Keep.
Frontend postinstallRuns npx playwright install.E2E test support.Build/install cost.Useful during validation; consider CI profile separation later.

1.2 Runtime Config Implications by Stage

Stage 0 — PR9.5, PR10, PR11, PR11.5

The main risk is validity contamination.

Stage 0 must freeze or explicitly record:

Model
Provider
Worker count
Reservoir preload state
Prompt set
Artifact set
Test conditions
Any memory/retrieval feature state

Stage 0 recommended settings posture:

RESERVOIR_LOAD_BOOT_DOCS=false
no vector retrieval
no graph memory
no scraper ingestion
fixed model/provider
fixed worker count
explicit max-run budget

Reason:

Stage 0 is proving artifact/claim mechanism.
Any hidden context, retrieval, or provider drift can invalidate the result.

Stage 1 — PR12–PR16

The main risk is promoting noisy signals.

Runtime behavior should still be conservative:

no automatic retrieval
no automatic claim extraction
no active reservoir scoring
no global heuristics
conversation_control treated as evidence, not authority

Settings to inspect if behavior is surprising:

AALAM_MODEL
provider keys / provider path
RESERVOIR_LOAD_BOOT_DOCS
QDRANT_URL
MEMORY_BACKEND_URL
rate-limit behavior

Stage 2 — PR17–PR21

The main risk is turning observed patterns into overconfident control.

Settings and config begin to matter more because validation rules, feedback, and confidence depend on stable execution.

Stage 2 PRs should record:

model/provider
validation rule version
artifact/claim set
feedback record format
confidence calculation method
any memory/retrieval state

Still avoid:

model self-confidence as truth
embedding-only context matching
hidden retrieval
unlogged provider fallback

Stage 3+ — PR22–PR43

Dormant infrastructure becomes progressively legitimate, but only after earlier gates pass.

Relevant settings become active later:

RESERVOIR_*        → Stage 3 memory / provenance / retrieval guardrails
QDRANT_*           → Stage 3 retrieval experiments, only after PR15/19
REDIS_*            → Stage 4–7 reliability, rate-limiting, queues
BOUNDARY_REGISTRY_*→ Stage 4 governed execution
WORKERS / pool     → Stage 7 performance scaling

Canonical rule:

A setting should not become active simply because it has a default.
It should become active because the current PR explicitly needs it.

1.3 Runtime Config Red Flags to Add to PR Review Checklist

Every PR through Stage 2 should include this quick check:

### Runtime / Cost / Validity Check

- [ ] Model/provider recorded.
- [ ] Total expected LLM calls recorded.
- [ ] Reservoir preload disabled or explicitly in scope.
- [ ] Vector/graph retrieval disabled or explicitly in scope.
- [ ] Scraper/ingestion disabled or explicitly in scope.
- [ ] Worker count known.
- [ ] No provider/model changed between comparison conditions.
- [ ] No hidden context source added.
- [ ] No default secret or wildcard CORS relied on for external use.

For Stage 6+ add:

- [ ] Non-default secrets enforced.
- [ ] Production CORS source of truth confirmed.
- [ ] Redis-backed or equivalent distributed rate limiting in place.
- [ ] Backend container non-root or exception documented.
- [ ] Forwarded proxy trust restricted.
- [ ] Endpoint auth/ownership audited.

1.4 Canonical Summary

Runtime config is not a side issue. In this project, runtime config can change the experiment.

The most important config risks are:

hidden reservoir preload
model/provider drift
weak/in-memory rate limits
default secrets
wildcard config defaults
worker-count multiplication
vector/graph memory activation
scraper or ingestion activation

Canonical Stage 0 rule:

Freeze runtime behavior before interpreting output behavior.

Section 2 — Endpoint / Auth Inventory

Purpose:

Separate routes that actually support the Stage 0 artifact loop
from routes that merely exist, appear protected, or belong to later-stage systems.

This section is not a full route-by-route security audit. It is a PR-drafting endpoint reality map: what future Aalams and engineers should rely on, ignore, inspect, or avoid activating.


2. Endpoint / Auth Reality

2.1 Key Endpoint Truths

The backend app includes a broad API router under /api, plus memory, Sub-AI, and drift routes. It also exposes GET /__acp/status, and the OpenAPI generator applies BearerAuth metadata globally to all paths.

That creates an important distinction:

OpenAPI says BearerAuth globally.
That does not prove every route enforces auth, ownership, or role checks.

For PR drafting, this matters because a PR should not say:

"endpoint is protected"

unless the route itself, dependency chain, middleware, or service layer proves protection.

Canonical endpoint rule:

Docs-level security metadata is not enforcement.
Enforcement must be verified in route/service behavior.

2.2 Endpoint Inventory by Stage Relevance

Endpoint / Route AreaCurrent RoleAuth / Ownership Reality to VerifyStage RelevancePR Drafting Guidance
Artifact create/list/getCore Stage 0 artifact loop.Earlier audit found static demo-owner behavior for artifact create/list/get, not real user ownership.Active now: PR9.5–PR21.Safe to use for validation. Do not claim production-grade ownership.
Aalam chat endpointsMain reasoning surface.Need route-level auth and cost controls verified before external use.Active now through PR43.Use as controlled reasoning surface; record model/provider and prompt conditions.
Memory routesIncluded separately under /api/v1/memory.Must verify whether memory is active, hidden, user-scoped, and auth-checked.Dormant / caution until Stage 3+.Do not rely on or activate for Stage 0–2 mechanism tests.
Sub-AI routesIncluded under /api/v1.Need provider/model/auth/cost audit.Later: PR33, PR39.Ignore until multi-agent/provider consistency or baseline comparison.
Drift monitoring routesIncluded under /api/v1/drift; drift middleware itself is commented out in main.py.Routes may exist even if scoring middleware is disabled.Diagnostic in PR13/16; later guardrail candidate.Treat as failure evidence / calibration candidate, not authority.
Governance routesExisting governance/ledger-style system from earlier audit.Must verify linkage to active artifact path.Later: PR17, PR27–31.Do not claim active governed execution until tied to validated rules/contracts.
Reservoir routesBroad content/reservoir CRUD from earlier audit.Must verify auth, business/user scoping, and whether boot preload is active.Dormant until PR14/15 reference; real use Stage 3+.Do not activate in Stage 0. Use as latent structure only.
System / health routesHealth/status and ACP status.May expose enforcement/status metadata.Useful for runtime validation and deployment.Useful in manifests, but not proof of artifact mechanism.
Admin / analytics / class / course / homework routesLegacy broader platform surface.Likely role-sensitive; not relevant to Stage 0.Later only if domain/workflow integration needs them.Ignore until Stage 5+ controlled integration.
Content generation / assessment routesPotential LLM-cost endpoints; rate limiter identifies these path patterns.Need endpoint-specific auth, run caps, rate limits, provider/model clarity.Mostly out of scope for Stage 0 unless explicitly used.Cost-sensitive; do not let validation accidentally call them.
Advanced memory routesRate limiter has a path case for /advanced-memory; hardcoded bypass appears inside that branch.Rate limiting may be disabled for this area.Dormant until memory stages.Do not use for Stage 0–2; audit before any external use.

2.3 Endpoint/Auth Drafting Checklist

Every PR through Stage 2 should answer:

### Endpoint / Auth Check

- [ ] Which endpoint(s) does this PR use?
- [ ] Is the endpoint part of the active Stage 0 artifact loop?
- [ ] Does the PR rely on real user ownership, or only static/demo owner behavior?
- [ ] Does the endpoint call an LLM?
- [ ] Does the endpoint access reservoir, memory, vector DB, scraper, or governance routes?
- [ ] Does the route actually enforce auth, or only appear protected in OpenAPI?
- [ ] Is there a cost/rate limit for this route?
- [ ] Is route behavior stable across repeated runs?

For Stage 6+ add:

- [ ] Route-level auth verified.
- [ ] Ownership checks verified.
- [ ] Role/override permissions verified if applicable.
- [ ] Anonymous access paths identified.
- [ ] OpenAPI security metadata compared against real enforcement.
- [ ] Abuse/rate-limit behavior tested.

2.4 Auth / Ownership Reality Map

Claim Future PR Might MakeSafe?Current Better Wording
“Artifacts are user-owned.”Not yet.“Artifacts are scoped through the current MVP/static owner path unless real ownership is verified.”
“Artifact reuse is authenticated and governed.”Not yet.“Artifact reuse is explicit and visible through Use in Chat; governance is not yet the active enforcement path.”
“Routes are protected because OpenAPI shows BearerAuth.”No.“OpenAPI advertises BearerAuth globally; route-level enforcement must be verified separately.”
“Memory/retrieval is available.”Too broad.“Memory/retrieval routes or infra exist, but Stage 0 uses explicit manual artifact reuse.”
“Drift monitoring protects outputs.”No.“Drift routes exist; drift middleware appears disabled, and prior behavior suggests calibration is needed.”
“Reservoir is the artifact system.”No.“Reservoir exists as a broader/dormant content system; the active Stage 0 mechanism is artifact create/list/get plus visible reuse.”

2.5 Endpoint Areas That Should Be Ignored Through Stage 2

Unless a PR explicitly scopes them, future Aalams should not draft Stage 0–2 PRs around:

advanced memory endpoints
reservoir CRUD as active retrieval
governance submit as active enforcement
admin dashboards
analytics endpoints
classroom/course/homework routes
audio/speech endpoints
cloud/storage endpoints
Sub-AI / multi-agent provider routes
scraper/crawler endpoints if any exist

Reason:

They may exist, but they are not needed to prove:
artifact → reuse → measurable improvement.

2.6 Endpoint Areas That Are Safe to Reference Now

Through Stage 0–2, the safest endpoint references are:

artifact create
artifact list
artifact detail
Aalam chat / Use in Chat surface
validation/test capture outputs
system health/status only for runtime context

Even here, wording should stay precise:

safe:
"uses current artifact endpoint"
"uses explicit visible reuse path"
"uses MVP/static owner behavior"

unsafe:
"uses production memory"
"uses governed retrieval"
"uses authenticated per-user artifact ownership"

2.7 Security/Cost Implications of Endpoint Reality

Endpoint behavior affects both security and cost.

Security implications

If auth is assumed from OpenAPI instead of verified in route code,
external validation may expose routes incorrectly.

If static/demo owner behavior is mistaken for real ownership,
artifact boundary claims will be overstated.

If reservoir/memory routes are active without ownership checks,
hidden cross-user or cross-context leakage risk increases.

Cost implications

Any endpoint that calls LLMs must be counted.

Any endpoint that retrieves boot docs, memory, reservoir content, or vector results
can increase prompt size and cost.

Any endpoint exempt from rate limiting can become a cost leak.

2.8 Canonical Summary

Endpoint reality for P9A:

The active mechanism is narrow:
artifact create/list/detail
+ visible Use in Chat
+ controlled validation captures.

Many other routes exist,
but they are either dormant, legacy, partial, or later-stage candidates.

Canonical PR drafting rule:

Do not draft a PR around an endpoint capability
unless the route behavior, auth boundary, cost behavior,
and stage relevance are verified.

Section 3 — LLM / Cost Call Inventory

Purpose:

Identify which repo structures can create LLM/API cost,
which settings can change cost,
and what every validation PR must record so cost and mechanism remain controlled.

This section is deliberately conservative. I did not find a clean route-level LLM call map from the repo search available here, so this section should be treated as a cost-control inventory from confirmed dependencies and runtime config, not a complete code-path audit.

Confirmed facts:

  • Backend dependencies include OpenAI, Anthropic, Google Generative AI, Dashscope, tiktoken, Qdrant, Redis, Graphiti/Neo4j, Selenium, audio/speech libraries, AWS SDK, and observability dependencies.
  • Backend settings include optional OpenAI / Anthropic / Google API keys and AALAM_MODEL defaulting to "gpt-4".
  • Backend Docker defaults to 4 Uvicorn workers unless overridden.
  • Rate limiting is in-memory, has high defaults, and contains a hardcoded bypass path for advanced memory.

3.1 LLM / Cost-Sensitive Inventory

Cost SourceConfirmed EvidenceWhy It Can Cost Money / ValidityStage Guidance
OpenAI provideropenai==1.91.0; OPENAI_API_KEY; AALAM_MODEL.Direct LLM generation cost. Model changes can invalidate comparisons.Allowed for Stage 0 only as fixed, recorded provider/model.
Anthropic provideranthropic==0.54.0; ANTHROPIC_API_KEY.Possible alternate provider cost/fallback.Must be disabled or explicitly recorded in validation runs.
Google Generative AIgoogle-generativeai==0.3.1; GOOGLE_API_KEY.Possible alternate provider cost/fallback.Same: no unrecorded provider changes.
Dashscopedashscope>=1.10.0.Additional LLM/provider path.Treat as dormant unless PR explicitly uses it.
AALAM_MODELDefaults to "gpt-4".Expensive default; also affects output quality and comparison validity.Every PR must record model; do not change model between conditions.
tiktokenPresent in requirements.Opportunity for token counting and budget control.Use for run manifests where practical.
Qdrant vector retrievalqdrant-client==1.14.2; QDRANT_URL.Retrieval can add long context and increase generation cost; can contaminate mechanism.Dormant through Stage 0–2 unless explicitly tested.
Graphiti / Neo4j memorygraphiti-core[neo4j]==0.18.9.Infra/runtime cost; hidden memory risk.Dormant until Stage 3 structured memory.
Redis / memory backendredis; REDIS_URL; MEMORY_BACKEND_URL.Infra cost; may alter memory/cache/rate behavior.Use deliberately; do not assume in-memory and Redis behavior match.
Reservoir preloadDefaults true via RESERVOIR_LOAD_BOOT_DOCS.Can load hidden context and increase prompt/context cost.Disable for Stage 0 validation.
Selenium / scraper stackSelenium, BeautifulSoup, lxml.Compute, network, storage, compliance cost.Dormant until controlled corpus ingestion.
Audio / speech stackEdge TTS, Azure speech, pydub, ffmpeg, SpeechRecognition.API/media processing cost.Ignore through Stage 2 unless audio PR is explicit.
AWS SDKboto3/botocore.Cloud storage/service costs.Dormant unless storage/cloud PR explicitly uses it.
PostHog frontend analyticsposthog-js.Analytics cost and privacy/compliance concerns.Do not use as validation evidence unless explicitly configured.
Playwright install/testsFrontend postinstall runs Playwright install; test scripts exist.CI/build/runtime cost.Acceptable for validation; separate CI profiles later if needed.
Multiple backend workers${WORKERS:-4}.Can multiply concurrent LLM calls and bypass effective rate limits.Fix worker count in validation manifests.
Rate limiterIn-memory; high limits; advanced memory branch bypass.Weak cost protection.Use explicit test-run caps; replace before external use.

3.2 Required LLM / Cost Record for Every Validation PR

Every PR from PR9.5 through PR21 that triggers model output should include:

### LLM / Cost Record

Provider:
Model:
Endpoint / code path:
Artifact count:
Prompt count:
Conditions per artifact:
Expected max generations:
Actual generations:
Approx token count, if available:
Reservoir preload active? yes/no
Vector retrieval active? yes/no
Graph memory active? yes/no
Scraper/ingestion active? yes/no
Audio/speech active? yes/no
Cloud/storage active? yes/no
Worker count:
Rate-limit / run cap:

Minimum acceptance rule:

If model/provider/conditions are not recorded,
the PR cannot make a reliable comparison claim.

3.3 Stage 0 Cost Budgets

Stage 0 PRs are validation PRs. They should have bounded run counts before execution.

PR9.5

Expected shape:

3–5 artifacts
× 5 conditions
= 15–25 core generations

Required conditions:

baseline
original
transformed
matched-prose substitute
degraded transformed

Cost-control instruction:

Do not add exploratory retries unless recorded separately.
Do not change model/provider mid-run.

PR10

Expected shape:

light volume, not open-ended stress

Before running PR10, define:

artifact count
prompt variants
negative cases
maximum total generations
stop condition

Cost-control instruction:

PR10 should not become an uncapped search for success cases.

PR11

PR11 is mostly infrastructure/evidence. It may not need many new model calls.

Cost-control instruction:

Prefer organizing existing PR7–10 evidence before generating new outputs.

PR11.5

Expected shape:

3–5 artifacts
× 3 conditions
= 9–15 core generations

Required conditions:

full artifact
claims-only
narrative-wrapper-only

Cost-control instruction:

If output is ambiguous, mark inconclusive rather than repeatedly regenerating until clear.

3.4 Cost Illusions / Invalid Comparisons

IllusionWhat HappensWhy It Invalidates PR EvidenceRequired Control
Model driftOne condition uses a different model/provider.Output difference may be model-caused, not artifact-caused.Same model/provider for all comparison conditions.
Retry fishingTester reruns until one output looks good.Biases evidence toward success.Record all generations or predefine rerun policy.
Hidden retrieval costReservoir/vector/memory adds unseen context.Increases cost and contaminates mechanism.Disable or explicitly record retrieval state.
Prompt-length driftTransformed artifacts are much longer than originals.Longer context may improve output independent of structure.Record input lengths or token counts where possible.
Parallel provider pathEndpoint may call fallback provider.Cost and behavior become hard to attribute.Confirm provider path in logs/config.
High worker countMultiple workers create concurrency.Can multiply cost and weaken rate limits.Fixed worker count for validation.
Analytics mistaken for evidenceProduct analytics show usage.Usage ≠ artifact-caused improvement.Use captured outputs and verdicts, not analytics alone.

Canonical rule:

A cost leak is also a validity leak
if it changes the context, model, or number of attempts.

3.5 Route / Code Path Cost Review Checklist

Before drafting or merging any PR that calls an LLM:

### Cost Path Review

- [ ] Which endpoint/code path calls the LLM?
- [ ] Which provider is used?
- [ ] Which model is used?
- [ ] Are there retries?
- [ ] Are there fallbacks?
- [ ] Is streaming used?
- [ ] Is token count logged?
- [ ] Is max output token length set?
- [ ] Is temperature or randomness controlled?
- [ ] Can reservoir/memory/vector context be added?
- [ ] Is rate limiting active for this endpoint?
- [ ] Is the route exempt from rate limits?
- [ ] Does worker count affect max concurrency?

For Stage 0–2, missing answers should not block every validation PR, but they should be recorded as limitations.


3.6 Recommended Cost-Safe Defaults for PR9.5–PR21

For controlled validation:

fixed provider
fixed model
fixed prompt set
fixed artifact set
fixed condition order
fixed worker count
reservoir preload off
vector retrieval off
graph memory off
scraper off
audio off
cloud side effects off
explicit max generations
all outputs captured

For PR review comments, use this phrasing:

This PR should not be evaluated until the model/provider,
generation count, and hidden-context state are recorded.

3.7 Canonical Summary

LLM/cost behavior is part of the experiment.

In P9A, cost control is not merely financial. It protects causal validity.

Same model.
Same provider.
Same conditions.
No hidden context.
No unrecorded retries.
No uncapped runs.

Canonical Stage 0–2 rule:

Every model call must be countable,
and every comparison must be attributable.

Section 4 — Dormant Feature Activation Map

Purpose:

Identify repo structures that exist but should not be treated as active capability
until the correct PR/stage explicitly activates them.

This section is especially important because the repos contain many systems that look useful: reservoir, memory, vector DB, graph memory, drift monitoring, scraper, analytics, audio, cloud, governance, and legacy education modules.

The core rule:

Existing code is not permission to use the code.
A feature becomes valid only when the PR gate requires it.

4.1 Dormant Feature Activation Table

Dormant / Latent FeatureCurrent Evidence / StateWhat Activates ItEarliest Safe StageWhy It MattersActivation Risk
Reservoir preloadRESERVOIR_LOAD_BOOT_DOCS defaults true; startup function can load boot docs into docking harness state.Runtime config enables preload and reservoir service returns boot docs.Stage 3+, unless explicitly tested earlier.Can add hidden context and contaminate artifact/claim mechanism tests.Stage 0 false positives: output improves because of boot docs, not artifact reuse.
Reservoir CRUD / content storeReservoir settings exist; prior audit found broad reservoir CRUD.PR explicitly scopes reservoir as storage/retrieval/memory structure.PR14–15 as reference only; real use Stage 3.Could support signal filtering, memory, provenance, authority later.Turns Stage 0/1 into generic CMS/retrieval project.
Vector retrieval / Qdrantqdrant-client dependency and Qdrant config exist.Code path queries Qdrant or injects vector-selected content into prompts.After PR15/19 manual matching evidence; generally Stage 3.Could later support retrieval behavior.Embedding similarity can hide invalid reuse and make mechanism opaque.
Graph memory / Neo4j / Graphitigraphiti-core[neo4j] dependency exists.Relationship graph is built or queried during reasoning/retrieval.Stage 3: PR22–26.Useful for artifact/claim relationship mapping.Graph-shaped noise if nodes/claims are not clean.
Redis memory/cache backendRedis dependency and memory backend settings exist.Runtime uses Redis for memory/cache/rate limits/queues.Rate limiting can matter earlier; memory use Stage 3+.Useful for distributed rate limits and reliability.Behavior may differ from local/in-memory tests.
Advanced memory endpointsRate limiter includes /advanced-memory path, with apparent hardcoded bypass.Requests hit advanced-memory routes.Stage 3+ after audit.May become memory/retrieval surface.Potential cost/security gap due to bypass.
Drift scoring middlewareDrift routes included; drift middleware is commented out as temporarily disabled.Middleware is uncommented or drift scoring used as decision input.PR13/16 as diagnostic; enforcement only later.Can become heuristic/guardrail candidate.Known risk of overblocking or false review flags.
Governance submit / ledgerGovernance concepts and ACP status exist; startup enforces boot manifest STRICT mode.PR connects governance route/ledger to active artifact/reuse decisions.Stage 2 for rule definition; Stage 4 for governed execution.Needed for contracts, trace, audit, override.Parallel governance path may be mistaken for active artifact mechanism.
Decision contracts / enforcement gateFoundation map says earlier PRs covered decision contract/enforcement concepts.Validated Stage 1–2 rules are converted into contracts/enforcement.Stage 4: PR27–31.Makes decisions enforceable.Enforcing unproven heuristics creates brittle failures.
Scraper / web parsing stackBeautifulSoup, lxml, Selenium exist.PR runs scraper/crawler or uses scraped content as artifacts/context.Controlled corpus in Stage 1+; broader use Stage 3/6.Could expand corpus and provenance tests.Network/compliance/cost risk; contaminates Stage 0 evidence.
File ingestion pipelineDependencies and prior audit found deterministic ingestion for limited file types.PR uses ingestion output as artifact source.PR12+ if corpus expansion is explicit.Useful for controlled corpus and provenance.Raw document context may be confused with artifact mechanism.
Audio / speech featuresAudio/speech dependencies exist: edge-tts, Azure speech, pydub, ffmpeg, SpeechRecognition.PR invokes speech/TTS/transcription/media paths.Not needed through Stage 2; domain-specific later.Could support language product eventually.API/media cost and irrelevant complexity.
Cloud / AWS servicesboto3/botocore exist.PR stores/reads artifacts or media from cloud.Later reliability/product stages unless explicitly scoped.Could support storage and deployment.Untracked cloud cost and secret exposure.
Frontend analytics / PostHogposthog-js dependency exists.Analytics events are enabled and used for evaluation.PR20+ only if feedback/evidence requires it.Could support external/user behavior analysis.Usage analytics ≠ mechanism evidence; privacy/compliance concerns.
Stripe / paymentsStripe dependencies exist in frontend.Payment/commercial flows are used.Out of scope until post-MVP/commercialization.Not relevant to artifact/reuse proof.Compliance and distraction.
Legacy classroom/course/homework modulesFrontend/backend platform includes broader education surfaces from prior audit.PR uses these workflows as domain surfaces.Stage 5+ integration; Stage 6 external validation.Useful as controlled domain workflows later.Product sprawl before core loop is proven.
Admin / dashboard surfacesBroad platform/admin/analytics surface exists.PR displays feedback, audit, confidence, or reliability dashboards.PR20 only for data model; UI later Stage 4/7.Could surface audit and reliability.Dashboard theater without real events.
Multi-agent / Sub-AI / provider pathsSub-AI route included; multiple provider dependencies exist.PR compares or coordinates multiple agents/providers.Stage 5 PR33; Stage 6 PR39.Useful for consistency and baseline tests.Provider differences can be mistaken for governance failure.

4.2 Dormant Feature Activation Rules by Stage

Stage 0 — PR9.5, PR10, PR11, PR11.5

Allowed active structures:

artifact create/list/detail
Use in Chat
manual artifact transformation
manual claim extraction
validation reports
Playwright captures
fixed Aalam model/provider

Must remain dormant:

reservoir preload
reservoir retrieval
Qdrant/vector search
Graphiti/Neo4j memory
advanced memory
scraper ingestion
automated claim extraction
governance enforcement
confidence scoring
analytics dashboards
audio/speech
cloud side effects

Reason:

Stage 0 is testing causal mechanism.
Dormant features add hidden variables.

Stage 1 — PR12–PR16

Allowed cautiously:

manual typology
manual claim extraction
failure taxonomy
signal filtering rubric
advisory heuristics
controlled corpus expansion if explicitly scoped

Still dormant:

automatic retrieval
vector matching
graph memory
global quality scoring
governance enforcement
automatic claim extraction
dashboarding

Reason:

Stage 1 identifies signal.
It does not yet automate signal.

Stage 2 — PR17–PR21

Allowed cautiously:

validation rules
reuse boundaries
manual/rule-based context matching
artifact-linked feedback records
offline confidence calibration

Still generally dormant:

hidden memory
automatic vector retrieval as authority
global enforcement
model self-confidence
dashboard-driven decisions
full reservoir activation

Reason:

Stage 2 controls reuse,
but control must be traceable to observed evidence.

Stage 3 — PR22–PR26

May activate, if prior gates pass:

relationship mapping
provenance tracking
consolidation
authority rubric
retrieval guardrails
reservoir as structured memory
graph/vector experiments
controlled ingestion

Activation condition:

Only after PR12–21 provide clean nodes, validated rules,
and known failure boundaries.

Stage 4 — PR27–PR31

May activate:

decision contracts
enforcement gate
traceability
audit events
human override
auth/role authority checks
governance ledger integration

Activation condition:

Only for rules proven in Stage 1–2.

Stage 5–7 — PR32–PR43

May activate:

cross-workflow modules
multi-agent/provider consistency
global constraints
degradation handling
external validation packets
adversarial testing
baseline comparisons
UX compression
performance/reliability infrastructure
MVP readiness definitions

Activation condition:

Only if core governed loop has survived Stage 0–4.

4.3 Activation Review Checklist

Every PR should include this question:

Did this PR activate any dormant feature?

If yes, require:

### Dormant Feature Activation Review

Feature activated:
Why this PR needs it:
Earlier proof this depends on:
Config/env change required:
Endpoint/code path affected:
Cost impact:
Security impact:
Validation impact:
Rollback plan:
Evidence that activation did not contaminate test:

If the PR cannot fill that out, the feature should remain dormant.


4.4 Common Accidental Activation Patterns

Accidental ActivationHow It HappensWhy It MattersPrevention
Reservoir preload silently activeDefault config remains true.Hidden boot docs may affect outputs.Set false and record state in manifest.
Vector retrieval accidentally includedQdrant path enabled by config or imported service.Output may reflect retrieved context, not artifact.Disable and log retrieval state.
Memory route used in test setupTester uses available memory endpoint for convenience.Breaks explicit reuse mechanism.Use artifact endpoints only through Stage 2.
Drift middleware re-enabledEngineer “fixes” review flags or monitoring.May block/alter outputs mid-validation.Treat as separate PR; do not combine.
Scraper used to enrich corpusEngineer adds source docs to make artifacts better.Corpus changes invalidate comparison.Freeze corpus per PR.
Automated claim extraction addedEngineer reduces manual work.Extraction errors become hidden system behavior.Manual extraction until claim unit is proven.
Analytics used as validationPostHog/user events collected.Usage is not output improvement.Use controlled outputs and verdict tables.
Dashboard added earlyEngineer makes audit/confidence visible.Looks mature before events are real.Build event record first, UI later.

4.5 Stage-Gated “Earliest Safe Use” Summary

Feature ClassEarliest Safe UseNotes
Artifact create/list/detailNowActive Stage 0 mechanism.
Use in ChatNowMust remain visible and user-controlled.
Manual transformed artifact formatPR9.5Validation only, not product schema.
Manual claim extractionPR9.5 / PR11.5Validation only, not automated.
Failure taxonomyPR13Do not fix before mapping.
Advisory heuristicsPR16Advisory only, not enforcement.
Validation rulesPR17Must cite Stage 1 evidence.
Reuse boundariesPR18Negative tests required.
Context matchingPR19Manual/rule-based first.
Feedback recordsPR20Artifact-linked, not generic thumbs.
ConfidencePR21Offline/calibrated only.
ReservoirPR14/15 as reference; Stage 3 activeDo not use for Stage 0.
Vector retrievalAfter PR15/19; usually Stage 3Never as proof by itself.
Graph memoryStage 3Clean nodes first.
Governance enforcementStage 4Proven rules only.
Audit/overrideStage 4Trace required.
Cross-workflow modulesStage 5Controlled integration only.
External usersStage 6Controlled first.
UX compressionStage 7Do not hide mechanism.

4.6 Canonical Summary

Dormant features are valuable because they reduce future build cost.

They are dangerous because they can invalidate current tests.

Latent infrastructure is an asset only if stage-gated.
Otherwise it becomes hidden behavior.

Canonical PR review rule:

If a PR activates reservoir, memory, vector retrieval, graph memory,
scraper ingestion, governance enforcement, analytics, audio, cloud,
or multi-agent/provider behavior,
the PR must explicitly say why that feature is now stage-appropriate.

Section 5 — Evidence Traceability Map

Purpose:

Define what counts as evidence,
where evidence should live,
and how future Aalams / engineers should determine whether a PR actually proved its claim.

This is the operational bridge between:

test ran

and:

claim proven

A PR through Stage 2 should not be considered complete merely because it contains screenshots, a narrative summary, or a passing test. It must produce evidence that can be inspected later.

Canonical rule:

A result is not evidence unless a future reviewer can reconstruct:
input → condition → output → comparison → verdict.

5.1 Evidence Types and Their Reliability

Evidence TypeUseful ForReliabilityMain GapRequired Pairing
Markdown validation reportHuman interpretation, verdict, failure explanation.MediumCan become narrative-only if not tied to captured outputs.Must link to artifact IDs, prompts, outputs, conditions.
Run manifestReproducibility and environment control.HighDoes not prove outcome by itself.Must pair with outputs and verdict table.
Captured text outputsActual model behavior.HighNeeds condition labels and comparison rubric.Must pair with prompt/artifact/condition metadata.
ScreenshotsUI state, visible reuse, evidence that flow happened.Low-to-mediumScreenshots alone do not prove mechanism.Must pair with extracted/captured text and verdict.
Playwright JSON / test artifactsRepeatability, browser automation, structured captures.High if completeCan be brittle or capture wrong selector.Must pair with human-readable verdict and artifacts.
LogsRuntime behavior, errors, hidden side effects, provider/model clues.MediumOften noisy and incomplete.Must pair with manifest and output evidence.
Code diffShows what changed.High for implementationDoes not prove behavioral effect.Must pair with tests/results.
Unit testsLocal correctness of functions.Medium-to-highMay not test end-to-end mechanism.Pair with e2e or validation output for mechanism PRs.
E2E testsUser-flow behavior.High for flowMay not prove causal artifact effect.Pair with controlled comparisons.
Human verdict tableMechanism classification.MediumSubjective if rubric unclear.Must include criteria and failure notes.
Analytics eventsUsage and behavior frequency.Low for mechanismUsage is not improvement.Only useful with artifact-linked outcomes.
Confidence scoreLater predictive calibration.Low until validatedCan be model theater.Must pair with observed outcomes.
Audit traceDecision reconstruction.High in Stage 4+Not available/needed for early PRs.Pair with contract/rule evidence.

5.2 Required Evidence Packet by PR Type

PR TypeRequired Evidence PacketMinimum Failure Evidence
PR9.5 transformationArtifact selection table; original artifact; transformed artifact; baseline/original/transformed/substitute/degraded outputs; verdict table.At least one case where transformed does not beat substitute or degradation does not weaken output, unless all cases pass.
PR10 stressRun manifest; artifact/prompt matrix; outputs for varied/ambiguous/degraded/wrong-artifact cases; failure-mode notes.At least one negative or degraded condition.
PR11 validation infrastructureStandard folder structure; manifest; input/output files; screenshots if applicable; verdict rubric; known limitations.Example of evidence that would be rejected as insufficient.
PR11.5 claim-level validationFull artifact, extracted claims, narrative-only wrapper, outputs for A/B/C, comparison table, mechanism classification.Case where A ≈ B ≈ C or classification is unclear.
PR12 pattern extractionCorpus table; artifact type; claim extractability; mechanism result; source verdict; pattern label.At least one artifact held/rejected rather than promoted.
PR13 failure mappingFailure taxonomy; example outputs; reproduction condition; likely cause; stage implication.At least one clear failure per major category available in corpus.
PR14 signal filteringSignal rubric; strong/weak examples; rejected artifact; rationale; counterexample.Artifact that sounds useful but is rejected.
PR15 selection behaviorTask → selected artifact → output → verdict table; wrong/plausible artifact cases.Plausible artifact that worsens or does not improve output.
PR16 heuristicsRule list; evidence source; A/B no-heuristic vs heuristic outputs; false-positive/negative examples.Heuristic failure case.
PR17 validation rulesRule table; source evidence; allowed/blocked action; test case; expected result.Rule violation example.
PR18 reuse boundariesBoundary table; allow/warn/block/review cases; negative test set.Wrong/stale/weak/mismatched artifact case.
PR19 context matchingMatching criteria; expected match/rejected match; output effect; mismatch examples.Plausible but wrong match.
PR20 feedback loopArtifact-linked feedback records; positive/negative/ambiguous outcomes; reviewer reason.Feedback that cannot be used because it is too generic.
PR21 confidenceConfidence bands; predicted outcomes; actual outcomes; mismatch table.High-confidence failure or low-confidence success.

5.3 Canonical Evidence Folder Structure

Recommended for validation PRs:

/docs/validation/prXX/
  manifest.md
  artifact_selection.md
  inputs.json
  outputs.json
  verdict_table.md
  failure_notes.md
  limitations.md
  screenshots/
  raw/

For PR9.5:

/docs/validation/pr9_5/
  manifest.md
  selected_artifacts.md
  transformed_artifacts.md
  comparison_outputs.json
  verdict_table.md
  aggregate_verdict.md
  failure_notes.md

For PR11.5:

/docs/validation/pr11_5/
  manifest.md
  selected_artifacts.md
  extracted_claims.md
  narrative_wrappers.md
  comparison_outputs.json
  mechanism_classification.md
  aggregate_summary.md

For PR12–16:

/docs/validation/stage1/
  corpus_table.md
  pattern_extraction.md
  failure_taxonomy.md
  signal_filtering.md
  selection_observations.md
  heuristic_tests.md

For PR17–21:

/docs/validation/stage2/
  validation_rules.md
  reuse_boundaries.md
  context_matching.md
  feedback_records.md
  confidence_calibration.md

5.4 Minimum Manifest Schema

Every validation PR should include:

# Manifest

PR:
Stage:
Branch / SHA:
Date:
Evaluator:
Model:
Provider:
Worker count:
Reservoir preload active: yes/no
Vector retrieval active: yes/no
Graph memory active: yes/no
Scraper/ingestion active: yes/no
Artifact set:
Prompt set:
Condition set:
Expected generation count:
Actual generation count:
Known limitations:

For later PRs add:

Validation rule version:
Reuse boundary version:
Context-matching method:
Feedback schema version:
Confidence method:
Governance contract version:
Trace format version:

5.5 Verdict Table Standard

A verdict should not simply say “better.”

Minimum verdict table:

CaseConditionOutput DifferenceMechanism ClaimVerdictFailure / Limitation
A1Artifact vs baselineMore specific feedback categoriesArtifact improved outputPASSBaseline already partially correct
A1Artifact vs substituteArtifact preserved exact rubric labelsStructure matteredPASSNeeds repeat across more artifacts
A6Artifact vs substituteSimilar learner guidanceContent, not structureFAIL / USEFUL CONTEXTNarrative artifact not constraint-bearing
A7Degraded vs originalNo meaningful weakeningStructure not load-bearingFAILDegradation may not have removed true claim

Canonical verdict labels:

PASS
FAIL
PARTIAL
INCONCLUSIVE
USEFUL_CONTEXT
MECHANISM_CONFIRMED
MECHANISM_NOT_PROVEN

Avoid:

good
better
strong
promising
seems useful
looks right

unless paired with explicit criteria.


5.6 Evidence Anti-Patterns

Anti-PatternWhy It FailsRequired Fix
Screenshot-only PRShows UI, not mechanism.Add captured text outputs and verdict table.
Narrative-only PRReviewer cannot reconstruct comparison.Add manifest, inputs, outputs, conditions.
One successful caseNo generality or negative test.Add baseline/substitute/degraded or failure case.
No artifact IDEvidence cannot be traced.Record artifact ID/source/body excerpt.
No model/providerComparison may be invalid.Record model/provider in manifest.
No condition labelsOutput cannot be interpreted.Label baseline/artifact/substitute/degraded.
No failure conditionPR cannot fail, so it cannot prove.Add explicit failure condition.
Unrecorded rerunsRetry fishing risk.Record all attempts or define rerun policy.
Model self-judges outputCircular evaluation.Use human/verdict rubric and captured outputs.
Analytics as proofUsage is not improvement.Link analytics only to artifact-level outcomes.

5.7 Evidence Traceability by Stage

Stage 0

Evidence must prove or disprove mechanism.

Required:

baseline
artifact/transformed/claims condition
negative condition
captured output
explicit comparison
human verdict
failure explanation

Stage 0 evidence is invalid if:

model/provider changed
hidden retrieval was active
artifact set changed mid-run
prompts changed between conditions
outputs were not stored

Stage 1

Evidence must classify signal and failure.

Required:

corpus table
artifact typology
failure taxonomy
signal rubric
rejected examples
source verdict links

Stage 1 evidence is invalid if:

all artifacts are promoted
failures are fixed before being mapped
typology is assumed rather than tested
heuristics are enforced before validation

Stage 2

Evidence must show controlled reuse.

Required:

validation rule source
boundary case
context-match case
feedback record
confidence/outcome comparison
negative tests

Stage 2 evidence is invalid if:

rules lack Stage 1 evidence
confidence is model self-report
feedback is generic thumbs-only
context matching is embedding-only without output validation

Stage 3+

Evidence must support memory, governance, integration, and product claims.

Required later:

provenance trace
relationship map
consolidation record
authority calibration
decision contract
enforcement trace
audit event
override record
external validation packet
baseline comparison
reliability replay

Stage 3+ evidence is invalid if:

graph/retrieval is assumed correct because it exists
authority is assigned without outcome history
enforcement blocks without explanation
external validation lacks controlled baseline
UX compression hides mechanism

5.8 Canonical Summary

Evidence is the system’s memory of what was actually proven.

For P9A:

No captured output → no evidence.
No comparison → no mechanism.
No negative test → no proof.
No manifest → no reproducibility.
No failure condition → no validation.

Canonical PR review rule:

A PR is not complete when the test passes.
A PR is complete when the evidence explains what passed,
what failed, what remains uncertain,
and what later PRs are now allowed to assume.

Section 6 — Active / Latent / Legacy Repo Reality Classification

Purpose:

Give future Aalams and engineers a fast way to decide:
- what can be used now
- what should wait
- what should be ignored
- what is security-sensitive
- what is cost-sensitive
- what requires additional audit before external use

This is a practical classification layer over the repo.

Canonical rule:

Do not ask “does this exist?”
Ask “is this active, validated, stage-appropriate, and safe to rely on?”

6.1 Repo Reality Classification

Structure / AreaCurrent ClassificationUse Now?Earliest Serious UseWhy
Artifact create/list/detailActive coreYesNowThis is the current Stage 0 storage/reuse substrate.
Use in Chat visible reuseActive coreYesNowThis is the current causal reuse surface.
Create Artifact UIActive coreYesNowSupports artifact creation for validation corpus.
Artifact panel/detail UIActive coreYesNowSupports inspection and manual reuse.
PR7/8/9 validation docsActive evidenceYesNowExisting evidence base for PR9.5–11.5.
Playwright capture specsActive evidence toolYesNowUseful for reproducibility and UI flow evidence.
Manual transformed artifact formatExperimental validation structureYes, manuallyPR9.5Tests whether narrative artifacts can become constraint-bearing.
Manual claim extractionExperimental validation structureYes, manuallyPR9.5 / PR11.5Tests whether claims are causal unit.
Rigid/narrative typologyProvisional findingYes, cautiouslyPR12Must be retested across broader corpus.
Conversation control / drift scoringProblematic latent structureNo, except diagnosticPR13 / PR16Use as failure evidence before heuristic authority.
Reservoir preloadLatent / dangerous for Stage 0NoStage 3+Can add hidden context and contaminate validation.
Reservoir CRUDLatent future memory/content storeNoStage 3, with limited reference in PR14/15Useful later, but scope-drift risk now.
Qdrant vector retrievalLatent future retrievalNoAfter PR15/19, usually Stage 3Retrieval must not precede manual matching proof.
Neo4j / Graphiti memoryLatent structured memoryNoStage 3Needs clean artifact/claim nodes first.
Redis memory/cacheLatent infraMaybe for rate limiting laterStage 4–7; rate limiting earlier if neededUseful for distributed runtime behavior, not mechanism proof.
Governance ledger / submit pathLatent governanceNo for Stage 0Stage 4, with rule work starting Stage 2Must be tied to validated rules/contracts.
Decision contracts / enforcement gateLatent governed executionNoStage 4Enforce only after rules are proven.
Audit / trace / override structuresLatent governance operationsNoStage 4Need contract and rule evidence first.
Scraper / ingestion pipelineLatent corpus/provenance systemNo for Stage 0PR12+ controlled; Stage 3 broaderCan contaminate mechanism if introduced too early.
Audio / speech modulesLegacy / domain featureNoDomain-specific laterNot relevant to artifact/reuse validation.
Admin / analytics dashboardsLegacy / later visibilityNoPR20 data first; UI laterDashboards without events are theater.
Classroom / course / homework modulesLegacy domain surfacesNoStage 5+May support cross-workflow tests later.
Sub-AI / multi-provider pathsLatent integrationNoPR33 / PR39Useful for consistency and baseline comparison later.
Stripe / payment featuresProduct/commercial legacyNoPost-MVPNot relevant to P9A validation.
PostHog analyticsLatent usage telemetryNo for evidencePR20+ cautiouslyUsage does not prove improvement.
Docker / deployment / health checksActive infraYes, for runtime contextNow; stronger use Stage 7Important for reproducibility, scaling, reliability.
Frontend Docker non-rootGood security patternYesNowUseful model for backend hardening.
Backend Docker root runtimeSecurity issueUse only internally with awarenessFix before external useRoot container is not product-ready.
OpenAPI BearerAuth metadataDocumentation signal onlyNo as proofAuth audit Stage 6+Does not prove route enforcement.

6.2 Active Now: Safe PR Building Blocks

These can be referenced directly in PR9.5–PR21.

artifact create/list/detail
visible Use in Chat
Create Artifact modal
Artifact panel/detail
manual validation reports
captured outputs
Playwright evidence
artifact corpus from PR8/PR9/PR10
manual claim extraction
manual transformed artifact format

Safe claim language:

The current system supports explicit, visible artifact reuse.
The current system supports storing and retrieving artifacts for controlled validation.
The current validation path can compare baseline/artifact/substitute/degraded outputs.

Unsafe claim language:

The system has governed retrieval.
The system has production memory.
The system has calibrated confidence.
The system has validated claim extraction.
The system has authenticated artifact ownership.
The system has product-ready governance.

6.3 Latent Later: Useful but Stage-Gated

These are probably valuable, but should not be activated until gates justify them.

reservoir
Qdrant
Neo4j / Graphiti
Redis-backed runtime state
governance ledger
decision contracts
enforcement gates
trace/audit/override
scraper ingestion
analytics
multi-provider / Sub-AI
cross-workflow legacy modules

Correct drafting posture:

Reference as existing structure that may be adapted later.
Do not make it the mechanism now.

Example:

Good:
“Reservoir may become the Stage 3 structured memory substrate if PR17–21 validate controlled reuse.”

Bad:
“Use reservoir retrieval in PR10 to improve stress-test results.”

6.4 Legacy / Ignore for Now

These are not necessarily bad. They are simply not relevant to the current proof.

classroom management
course/homework modules
audio/speech product features
payment/Stripe
broad analytics dashboards
admin surfaces
gamification / multimedia surfaces
general scraper/web-crawling expansion

Ignore rule:

If a module does not help prove artifact → reuse → measurable improvement,
ignore it through Stage 2.

Later, some of these may become useful for:

Stage 5 cross-workflow integration
Stage 6 external validation
Stage 7 UX/product readiness

6.5 Security-Sensitive Structures

These deserve special handling before external users or product readiness.

StructureRiskWhen to Fix / Audit
Backend container running as rootContainer compromise blast radius.Before Stage 6 external validation, definitely before Stage 7.
Default secretsUnsafe if env missing.Add production startup check before external users.
CORS split between allowlist and wildcard configFuture misconfiguration risk.Consolidate before external users.
--forwarded-allow-ips "*"Proxy header spoofing risk.Restrict before external deployment hardening.
OpenAPI global BearerAuthMay imply protection without enforcement.Endpoint auth audit before external use.
Static/demo owner artifact behaviorNot real ownership.Replace before real users / multi-user tests.
In-memory rate limiterWeak across workers/instances.Replace with Redis-backed limits before external use.
Advanced-memory rate-limit bypassPotential abuse/cost gap.Audit before memory activation.
Reservoir/memory routesPotential hidden context / cross-context leakage.Audit before Stage 3+ activation.
Scraper/cloud/audio integrationsExternal data/API risk.Audit before activation.

6.6 Cost-Sensitive Structures

These can materially affect spend or invalidate controlled tests.

StructureCost RiskStage Guidance
Aalam LLM callsDirect model cost.Record provider/model/calls every PR.
Multiple providersHidden fallback/parallel costs.Disable or record provider path.
AALAM_MODEL="gpt-4" defaultHigher generation cost.Use deliberately; do not change mid-test.
Reservoir preloadMore context, hidden retrieval, possible extra processing.Disable Stage 0.
Vector/graph retrievalInfra + larger prompts.Stage-gated.
Scraper/SeleniumCompute/network/storage cost.Controlled use only.
Audio/speechAPI/media processing cost.Out of scope through Stage 2.
AWS/cloudStorage/service costs.Out of scope unless explicit.
Playwright postinstall/testsCI/build time.Acceptable for validation; optimize later.
Backend workersMultiplies concurrency and rate-limit gaps.Fix worker count in validation.
Analytics/PostHogUsage data cost/compliance.Not validation evidence by itself.

6.7 Test-Sensitive Structures

These can change the interpretation of results even if they are not “features.”

StructureWhy Test-SensitiveControl Needed
Model/providerChanges output quality.Fixed across conditions.
Prompt wordingCan cause apparent improvement.Same prompt across comparison conditions.
Artifact body editsCan create success by rewriting.Freeze artifact per condition.
Transformed artifact lengthLonger text may improve output independent of structure.Record input length/token count where practical.
Reservoir preloadHidden context.Disable or record explicitly.
Drift/conversation controlMay block or alter outputs.Record state; do not treat as authority.
Worker countCan affect concurrency/rate behavior.Record in manifest.
Retrieval/memoryAdds hidden variables.Disable unless explicitly in scope.
RerunsCan bias toward good output.Record all attempts or define rerun policy.

6.8 PR Review Questions: Active vs Latent vs Legacy

Every future PR should answer:

### Repo Reality Review

1. Which active Stage-appropriate structures does this PR use?
2. Which latent structures does this PR intentionally avoid?
3. Did this PR activate any dormant feature?
4. Did this PR rely on a UI surface that is not backed by verified behavior?
5. Did this PR rely on static/demo behavior as if it were production behavior?
6. Did this PR touch legacy modules unnecessarily?
7. Did this PR create new cost/security exposure?
8. Did this PR preserve comparison validity?

If the answer to #3 is yes, require the Dormant Feature Activation Review from Section 4.


6.9 Canonical Summary

Repo reality classification:

Active now:
artifact storage + visible reuse + validation evidence.

Latent later:
reservoir + retrieval + memory + governance + audit + scraper + analytics.

Legacy / ignore now:
broad education product, audio, payment, admin dashboards, domain sprawl.

Security-sensitive:
auth, secrets, CORS, root container, rate limits, owner boundaries.

Cost-sensitive:
LLM calls, providers, retrieval, preload, scraper, audio, cloud, workers.

Canonical engineering rule:

Use the narrow active mechanism until the evidence earns the right to activate the broader system.

Section 7 — Practical Review Packets for Engineers

Purpose:

Define the concrete packets an engineer or future Aalam should pull
when reviewing BE, FE, scraper, infra, or validation PRs.

This section converts the prior review into practical review checklists.

Canonical rule:

A good repo review produces artifacts engineers can use directly in PR comments,
acceptance criteria, and go/no-go decisions.

7.1 Backend Review Packet

A backend review should pull:

ArtifactWhat to inspectWhy
Endpoint mapRoutes, methods, auth dependencies, ownership checks, LLM calls.Prevents relying on routes that only appear protected or active.
Config mapEnv vars, defaults, feature flags, model/provider settings.Prevents hidden behavior, cost drift, and config contamination.
Data model mapArtifactVersion, OwnerVault, reservoir, claims/provenance if present.Clarifies what can safely be extended.
Service dependency mapPostgres, Redis, Qdrant, Neo4j, reservoir, boundary registry, cloud APIs.Shows what must run and what can fail.
Cost path mapLLM calls, retries, provider fallback, token limits, model selection.Controls spend and comparison validity.
Security mapSecrets, auth, CORS, container user, rate limits, ownership.Identifies production blockers.
Failure behavior mapTimeouts, fallback behavior, error handling, fail-open/fail-closed.Prevents silent degradation.
Test mapUnit, integration, e2e, validation reports.Shows what behavior is protected.

Backend review should answer:

1. Which backend path is active for this PR?
2. Does it call an LLM or external service?
3. Does it rely on static/demo ownership?
4. Does it activate reservoir, memory, vector retrieval, or governance?
5. Does it fail closed or silently continue?
6. Is cost bounded?
7. Is endpoint auth real or only documented?
8. What evidence proves the backend behavior?

7.2 Frontend Review Packet

A frontend review should pull:

ArtifactWhat to inspectWhy
User-flow mapCreate Artifact, artifact list/detail, Use in Chat, chat send.Confirms what the user actually does.
State-flow mapStores, pending input, selected artifact, message state.Identifies hidden or visible handoffs.
API-client mapWhich backend endpoints the UI calls.Prevents UI assumptions about backend behavior.
Evidence-capture mapPlaywright specs, screenshots, JSON outputs.Shows what tests actually prove.
UX illusion mapUI surfaces that look more capable than they are.Prevents overclaiming retrieval, memory, feedback, confidence.
Cost trigger mapButtons/actions that trigger LLM calls.Prevents accidental expensive flows.
Error-state mapLoading, failure, no-response, blocked output states.Important for PR10/13/35.
Analytics mapPostHog or event tracking, if used.Ensures analytics are not mistaken for mechanism evidence.

Frontend review should answer:

1. Is reuse visible to the user?
2. Does the UI hide or inject artifact content?
3. Does this PR change the test flow?
4. Does the UI imply retrieval/ranking/confidence that does not exist?
5. Are outputs captured as text, not only screenshots?
6. Are error/loading states recorded?
7. Does the frontend call the expected backend endpoint?
8. Did UX changes alter validation comparability?

7.3 Scraper / Ingestion Review Packet

A scraper or ingestion review should pull:

ArtifactWhat to inspectWhy
Source inventoryWhat sources/files are ingested.Controls corpus and provenance.
Extraction supportFile types, parser behavior, PDF/DOCX gaps.Prevents assuming unsupported extraction.
Provenance mapSource path, timestamp, actor, version hash.Required for Stage 3 trust and lineage.
Dedupe behaviorHash-based, semantic, or none.Clarifies duplicate handling.
Chunking behaviorWhether content is chunked, embedded, indexed.Determines retrieval/memory implications.
Network behaviorWhether scraper hits external sites.Security/cost/compliance risk.
Storage behaviorRaw file, extracted text, artifact version, metadata.Defines what can be audited later.
Failure behaviorParser failure, partial extraction, unsupported file.Prevents silent bad corpus.

Scraper review should answer:

1. Is this controlled ingestion or open scraping?
2. What exact source became what artifact/text?
3. Is provenance preserved?
4. Is extraction complete or partial?
5. Are unsupported files clearly marked?
6. Is content deduped?
7. Is anything embedded or retrieved automatically?
8. Could ingestion contaminate artifact-mechanism tests?

Through Stage 0:

Scraper / ingestion should generally remain off.

Earliest safe controlled use:

PR12+ for corpus expansion,
Stage 3 for provenance / structured memory,
Stage 6 for external validation if tightly controlled.

7.4 Infra / Deployment Review Packet

An infra review should pull:

ArtifactWhat to inspectWhy
DockerfilesBase images, root/non-root user, installed packages.Security and reproducibility.
Runtime commandWorkers, proxy headers, host/port.Concurrency, security, cost.
Health checksWhat is checked and what is not.Health ≠ mechanism.
Env managementSecrets, defaults, missing validation.Prevents unsafe production boot.
CI/CD scriptsBuild/test/deploy commands, root SSH, hard resets.Operational risk.
Service dependenciesDB, Redis, Qdrant, Neo4j, reservoir, boundary registry.Runtime failure modes.
Logs / observabilityWhat is logged, where, and with what sensitivity.Debugging and audit.
Scaling behaviorWorkers, connection pool, queues, rate limits.Cost and reliability.

Infra review should answer:

1. Does deployment run as non-root?
2. Are default secrets rejected?
3. Are proxy headers trusted safely?
4. Are rate limits distributed?
5. Are health checks meaningful?
6. Does worker count change cost or test validity?
7. Are dormant services running unnecessarily?
8. Can a validation run be reproduced in the same environment?

7.5 Validation PR Review Packet

A validation PR review should pull:

ArtifactWhat to inspectWhy
ManifestModel, provider, env, artifact set, prompt set, conditions.Reproducibility.
InputsArtifact body, claims, transformed version, prompt.Comparison integrity.
OutputsCaptured model responses.Actual behavior.
Negative testsSubstitute, degraded, wrong artifact, narrative-only, etc.Prevents false positives.
Verdict tableExplicit pass/fail/inconclusive classification.Prevents subjective drift.
Failure notesWhat failed and why.Feeds future PRs.
LimitationsKnown gaps and uncertainty.Prevents overclaiming.
Run budgetExpected and actual generations.Cost control.

Validation PR review should answer:

1. What exact claim did this PR test?
2. What would have counted as failure?
3. Did it include baseline?
4. Did it include negative/degraded/substitute condition?
5. Were all outputs stored?
6. Did model/provider stay fixed?
7. Was hidden retrieval/memory disabled?
8. Does the verdict justify the next PR dependency?

7.6 Security Review Packet

A security review should pull:

ArtifactWhat to inspectWhy
Secrets/defaultsSECRET_KEY, JWT_SECRET_KEY, API keys.Prevents unsafe boot.
Auth enforcementRoute dependencies, middleware, ownership checks.Prevents overclaiming protection.
CORS configActive origins vs config defaults.Prevents browser exposure.
Container userRoot vs non-root.Reduces blast radius.
Rate limitingIn-memory vs distributed, bypasses.Abuse and cost control.
External integrationsLLMs, scraper, cloud, analytics, audio.Data/cost/compliance risk.
LoggingSecrets or sensitive content in logs.Data exposure risk.
Override pathsHuman override / governance bypass.Ensures traceability.

Security review should answer:

1. Are default secrets impossible in production?
2. Are route protections real?
3. Are ownership boundaries real?
4. Can untrusted headers affect identity/rate limiting?
5. Can expensive routes be abused?
6. Can hidden memory leak context?
7. Are external services stage-appropriate?
8. Are overrides traced?

7.7 Cost Review Packet

A cost review should pull:

ArtifactWhat to inspectWhy
LLM call mapProvider, model, endpoint, trigger, retries.Direct cost.
Run budgetExpected vs actual generations.Prevents runaway validation.
Context sourcesArtifact, reservoir, memory, retrieval, scraper.Token cost and validity.
Worker/concurrency configUvicorn workers, queues, frontend tests.Multiplies cost.
Rate limitsActive limits and bypasses.Cost guardrail.
External APIsSpeech, cloud, analytics, scraper.Non-LLM costs.
CI/build costPlaywright install, heavy deps.Operational cost.

Cost review should answer:

1. What action triggers model/API calls?
2. How many calls can this PR make?
3. Is there a hard cap?
4. Are retries/fallbacks recorded?
5. Does context size change across conditions?
6. Are dormant services adding hidden cost?
7. Are rate limits real under multiple workers?
8. Is the cost worth the PR claim being tested?

7.8 Go / No-Go Review Summary

For a PR through Stage 2, the reviewer can use this compact rule.

GO

The PR uses active stage-appropriate structures.
The runtime state is recorded.
The cost is bounded.
The endpoint behavior is verified.
The evidence is reproducible.
The failure condition is explicit.
The PR does not activate dormant systems.

NO-GO

The PR relies on UI illusion.
The PR changes model/provider mid-comparison.
The PR activates hidden retrieval/memory.
The PR has no negative test.
The PR has screenshots but no captured outputs.
The PR uses generic “better” judgments.
The PR has no run budget.
The PR cannot say what would count as failure.

CONDITIONAL GO

The PR may proceed if the limitation is explicitly recorded
and the next PR is not allowed to assume more than the evidence proves.

7.9 Canonical Summary

A thorough repo review should not merely produce comments like:

looks good
needs tests
security concern
cost concern

It should produce review packets that let future engineers know:

what path is active
what path is dormant
what evidence exists
what was not proven
what setting could change behavior
what cost/security risk exists
what later PR is now allowed to assume

Canonical final rule for engineers:

Review the repo as it actually runs,
not as the UI suggests,
not as the architecture hopes,
and not as the roadmap expects.

Section 8 — Canonical Repo Reality Review Summary

Purpose:

Condense the operational review into the small set of rules future Aalams and engineers should carry forward.

This section closes the canonical document.


8.1 The Core Repo Reality

The repo is not empty, immature, or unusable.

It is better described as:

A broad existing FE / BE / scraper / infra codebase
with a narrow active Stage 0 mechanism
and many latent structures that may become useful later.

The active mechanism today is:

artifact create/list/detail
→ visible Use in Chat
→ controlled model response
→ validation evidence

The latent system includes:

reservoir
memory
vector retrieval
graph memory
scraper/ingestion
governance ledger
decision contracts
enforcement
audit
override
analytics
multi-agent/provider paths
legacy domain workflows

Those latent structures are valuable, but only if stage-gated.

Canonical interpretation:

The repo is ahead of the proof.
That is useful only if the proof remains in control.

8.2 What Can Be Confirmed

Yes:

Most existing FE / BE / scraper structures can be used and adapted
IF each PR/stage gate passes
AND IF each structure is activated only at the appropriate stage.

This means:

Structure ClassCurrent StatusUse Posture
Artifact storage/reuse UIActive nowUse directly for Stage 0–2 validation.
Validation docs/capturesActive nowStrengthen into reproducible evidence.
Manual claim/artifact transformationActive as validation methodUse manually before schema/product changes.
Reservoir/vector/graph/memoryLatentHold until signal/control evidence justifies activation.
Governance/enforcement/audit/overrideLatentHold until validated rules exist.
Scraper/ingestionLatentControlled corpus only after mechanism proof.
Legacy product modulesLater-stage surfacesIgnore until cross-workflow/domain testing.
Infra/deployment/securityReal and importantHarden progressively before external/product stages.

8.3 Most Important Risks

The main risks are not that the repo lacks features.

The main risks are:

using too many features too early
mistaking UI surfaces for mechanism
mistaking stored text for governed artifacts
mistaking useful context for constraint-bearing structure
mistaking model quality for artifact effect
mistaking screenshots for evidence
mistaking config defaults for safe runtime behavior

The highest-risk false positives:

False PositiveWhy It Matters
Output improves because wording improved, not artifact structure.PR9.5 false success.
Claims-only performs well because it is shorter/clearer, not because claims are causal.PR11.5 misread.
Reservoir/vector/memory adds hidden context.Invalidates Stage 0 mechanism tests.
UI shows artifact panel, but actual reuse is manual text handoff.Overclaims retrieval/memory.
OpenAPI shows BearerAuth, but route enforcement is not verified.Security overclaim.
Confidence/ranking fields exist but are not calibrated.False authority.
Feedback exists but is not artifact-linked.No real learning loop.

8.4 Security Conclusions

The current security posture is acceptable for controlled internal validation if carefully configured, but it is not enough for external/product use without hardening.

Most important security flags:

backend container currently runs as root
default secrets exist in config
CORS has mixed source of truth
forwarded IP trust is broad
rate limiting is in-memory and partly bypassed
OpenAPI security metadata does not prove real route auth
static/demo artifact ownership is not production ownership
reservoir/memory routes need auth/ownership audit before activation

Stage guidance:

Stage 0–2:
  record and constrain security-relevant runtime behavior

Stage 6+:
  harden secrets, CORS, rate limiting, container user, route auth, ownership

Stage 7:
  reliability/security must be part of MVP readiness

8.5 Cost Conclusions

Cost control is part of validation integrity.

Most important cost-sensitive structures:

AALAM_MODEL default
provider/model selection
number of generations
reruns
worker count
reservoir preload
vector/graph retrieval
scraper/Selenium
audio/speech
cloud services
analytics
rate-limit bypasses

Canonical validation cost rule:

Every model call must be countable.
Every comparison must be attributable.

For PR9.5–PR21, every PR should record:

provider
model
artifact count
prompt count
condition count
expected generations
actual generations
hidden-context state
worker count
run cap

8.6 Endpoint/Auth Conclusions

Endpoint reality:

Only the narrow artifact + chat path should be assumed active for Stage 0.

Other routes may exist,
but they should not be treated as validated capability
until their auth, ownership, cost, stage relevance, and behavior are verified.

Safe wording:

The current system supports explicit visible artifact reuse.

Unsafe wording:

The current system supports governed retrieval, production memory,
calibrated confidence, or authenticated artifact ownership.

8.7 Evidence Conclusions

The central evidence rule:

No captured output → no evidence.
No comparison → no mechanism.
No negative test → no proof.
No manifest → no reproducibility.
No failure condition → no validation.

Every validation PR should make it possible to reconstruct:

input
→ condition
→ output
→ comparison
→ verdict
→ implication for next PR

A passing test is not enough.

A PR is complete only when it states:

what passed
what failed
what remains uncertain
what later PRs may now assume
what later PRs may not assume

8.8 Stage-Gated Repo Use Summary

StageRepo Structures to UseStructures to Avoid
Stage 0Artifact endpoints, Use in Chat, validation captures, manual transformations/claims.Reservoir, vector/graph memory, scraper, governance enforcement, schema/product changes.
Stage 1Typology, failure taxonomy, signal rubric, manual selection observations.Retrieval automation, global scoring, enforcement, dashboards.
Stage 2Validation rules, reuse boundaries, context matching, feedback records, offline confidence.Hidden memory, embedding-only authority, model self-confidence.
Stage 3Provenance, relationships, consolidation, authority, guardrails, cautious reservoir/memory.Graph/vector authority before clean nodes.
Stage 4Decision contracts, enforcement, trace, audit, override.Generic governance not tied to evidence.
Stage 5Cross-workflow, multi-agent consistency, global constraints, degradation handling.Product sprawl.
Stage 6Controlled external users, adversarial tests, baseline comparison.Open beta / marketing claims.
Stage 7UX compression, performance, reliability, MVP readiness.New feature expansion that changes what was proven.

8.9 Final Engineering Checklist

Every PR should include this compact review block:

## Repo Reality / Validation Check

- [ ] Active structures used:
- [ ] Dormant structures avoided:
- [ ] Dormant structures activated, if any:
- [ ] Model/provider recorded:
- [ ] Expected and actual generation count:
- [ ] Reservoir/vector/graph/scraper state recorded:
- [ ] Endpoint/auth assumptions verified:
- [ ] Static/demo behavior not overclaimed:
- [ ] Evidence packet included:
- [ ] Negative/failure condition included:
- [ ] Later-stage assumption allowed by this PR:
- [ ] Later-stage assumption NOT allowed by this PR:

For any dormant feature activation:

## Dormant Feature Activation Review

Feature:
Why now:
Earlier proof:
Config change:
Endpoint/code path:
Cost impact:
Security impact:
Validation impact:
Rollback:
Evidence it did not contaminate test:

8.10 Final Canonical Rules

1. Use the narrow active mechanism until evidence earns broader activation.
2. A repo feature should not move forward because it exists.
It should move forward only when the current PR gate needs it.
3. Validation first.
Schema second.
Product behavior last.
4. Do not make the system look smarter by making the experiment less controlled.
5. A later PR may only enforce what an earlier PR has proven.
6. Every PR must prove the smallest claim it needs,
with the least system change possible,
and with at least one way it could fail.
7. Continue only if reuse produces observable, attributable improvement.

That completes the canonical repo reality review.

The full document now has eight sections:

1. Purpose, Scope, and Runtime Config Inventory
2. Endpoint / Auth Inventory
3. LLM / Cost Call Inventory
4. Dormant Feature Activation Map
5. Evidence Traceability Map
6. Active / Latent / Legacy Repo Reality Classification
7. Practical Review Packets for Engineers
8. Canonical Repo Reality Review Summary

Final consolidation statement:

Most existing FE / BE / scraper / infra structures are reusable,
but not all are active, validated, or stage-appropriate today.

The active Stage 0 system is narrow:
artifact storage
+ visible reuse
+ controlled validation evidence.

Everything else should be treated as latent until a PR gate proves it is needed.

The canonical working rule for Abhishek / future Aalams:

Do not ask:
“What can the repo do?”

Ask:
“What has the repo proven,
under controlled evidence,
at this stage of the PR sequence?”

No further section is needed unless you want this converted into a GitHub issue / PR review template / .md handoff.