Project Overview
Artificial intelligence systems are increasingly deployed in high-stakes institutional environments—medicine, public administration, regulatory compliance, and large organizations—yet they remain structurally unsuited to those contexts. Most contemporary AI systems are designed as stateless or weakly stateful tools whose outputs are fluent but ephemeral, optimized for interaction rather than accountability. Their “memory,” where it exists, is typically limited to personalization or retrieval optimization and is deliberately decoupled from responsibility due to legal and ethical risk. As a result, institutions face a paradox: AI systems are used to inform decisions, but they are not designed to own or explain the institutional memory those decisions depend upon.
This project proposes and evaluates Agora-AI, an audit-first AI governance and memory framework designed to support institutional decision-making without requiring trust in model correctness. Rather than attempting to make models inherently safe or universally accurate, Agora-AI constrains how models are used by embedding them within a governed structure that preserves provenance, rationale, uncertainty, and reviewability. The system treats AI outputs not as transient answers, but as auditable decision artifacts that persist, can be contested, and can be evaluated over time.
The central hypothesis of this project is that institutional risk from AI systems arises less from model error than from the absence of structured memory and governance. When decisions cannot be reconstructed, responsibility dissolves, errors propagate silently, and compliance becomes performative rather than substantive. By contrast, when AI contributions are captured as structured claims with explicit evidence, permissions, and review paths, institutions can tolerate model fallibility while preserving accountability.
Agora-AI operationalizes this hypothesis by introducing a governed orchestration layer between AI models and institutional use. This layer enforces role-based access to information, logs all contextual inputs and transformations, constrains how outputs may be promoted into operational decisions, and preserves a durable memory of decisions and their rationales. The result is not an autonomous decision system, but an institutional memory and governance substrate capable of hosting AI assistance in domains where auditability is essential.
The project will focus primarily on a medical use case, where longitudinal context, regulatory scrutiny, and decision traceability are unavoidable. A secondary focus will examine institutional management workflows, where policy drift and loss of rationale are persistent problems. Law and democratic governance are treated as adjacent domains whose requirements inform design constraints but are not the primary evaluation targets in this phase.
Specific Aims
Aim 1: Design and implement an audit-first AI memory and governance architecture suitable for regulated institutional environments.
The first aim is to design and implement the core Agora-AI architecture as a reusable, domain-agnostic framework. This architecture will include: an ingestion layer for approved institutional artifacts with provenance tagging; a claim layer that forces AI outputs into structured assertions with uncertainty and evidence; a governance layer that enforces role-based permissions and promotion gates; a memory layer that preserves decision artifacts and their dependencies over time; and an observability layer that enables retrospective audit of all AI-assisted decisions.
Success for this aim will be defined by the system’s ability to reproduce, for any given decision, what information was used, what the AI asserted, what uncertainties were present, and what human approvals or overrides occurred. The architecture will be explicitly designed to be inspectable and regulator-friendly, prioritizing legibility over optimization.
Aim 2: Deploy and evaluate Agora-AI in a constrained medical decision-support workflow.
The second aim is to deploy Agora-AI in a narrowly defined medical context where decision provenance and longitudinal memory are critical, such as guideline-based chronic disease management or medication reconciliation. The system will not replace clinical judgment; instead, it will mediate how AI-generated insights are presented, recorded, and reviewed.
Evaluation will focus on institutional outcomes rather than model performance. Metrics will include the completeness of documented decision rationale, the frequency with which conflicting information is surfaced by the system, the time required to identify and correct errors, and the effort required for retrospective audit compared to baseline workflows. This aim tests whether an audit-first memory layer improves safety and accountability even when AI outputs are imperfect.
Aim 3: Measure the impact of governed AI memory on institutional drift and accountability in organizational management contexts.
The third aim extends evaluation to institutional management, where failures often arise not from lack of policy but from erosion of rationale and undocumented exceptions. Agora-AI will be applied to internal decision processes such as policy enforcement, risk review, or operational exceptions.
The evaluation will measure whether the system can detect patterns where informal practice diverges from formal policy, and whether those divergences become visible early enough to be addressed. Success will be defined not by enforcement, but by the system’s ability to surface latent structural inconsistencies in a way that is actionable rather than punitive.
Aim 4: Assess privacy, ethical, and regulatory implications of audit-first AI memory.
The final aim is to assess whether Agora-AI’s design mitigates common ethical risks associated with AI in institutions, particularly surveillance, over-collection of data, and opacity. The project will evaluate whether role-based memory, purpose limitation, and explicit governance gates can support compliance with medical and organizational privacy norms while still enabling meaningful auditability.
This aim will produce a set of design guidelines and policy recommendations describing how AI systems can retain institutional memory without becoming instruments of coercive monitoring or liability amplification.
Expected Contributions
This project will contribute a tested architectural model for AI governance grounded in institutional reality rather than consumer interaction paradigms. It will provide empirical evidence that accountability in AI systems can be improved through structure rather than control, and that meaningful governance is compatible with, and even enhanced by, model fallibility when memory and decision artifacts are made auditable.
By reframing AI safety and governance as problems of institutional memory and structure, Agora-AI offers a path toward responsible AI deployment in medicine and beyond—one that is measurable, reviewable, and compatible with existing regulatory frameworks rather than opposed to them.
Member discussion: