Abstract
This paper presents a case study contrasting an ACP-governed AI interaction with typical enterprise AI workflows in the production of rapid-turn policy intelligence for U.S. government (USG) personnel. Using a live sequence of loosely specified user requests—spanning web search, source differentiation, quote extraction, synthesis, and packaging—we examine how governance emerges (or fails) across turns. The study demonstrates that while enterprise AI can replicate surface capabilities (search, summarization, formatting), it routinely collapses provenance, smooths uncertainty, and launder responsibility under time pressure. By contrast, the ACP approach foregrounds authority clarity, failure-as-output, and inspection survivability. The interaction culminates in a compressed, attribution-safe briefing packet that remains usable without overclaiming. We conclude that prompt engineering alone cannot enforce governance at scale; enforceable schemas, tool constraints, and human interruption rights are necessary to prevent institutional risk.
1. Introduction
Organizations increasingly rely on AI systems to assemble policy-relevant information under tight timelines. Enterprise AI tools promise speed and synthesis, but often lack mechanisms to preserve provenance, signal uncertainty, or refuse unsafe compression. This paper analyzes a real interaction in which an analyst requested loosely defined research on U.S.–India trade developments, embassy messaging, and senior official remarks. The absence of explicit format or attribution requirements mirrors real-world conditions. The question is not whether AI can produce an answer, but whether it can produce an institutionally survivable artifact.
2. Case Description
2.1 Initial Conditions
- User intent: Situational awareness and quick-turn briefing for USG personnel.
- Constraints: Public, unclassified sources; attribution required; rapid turnaround.
- Prompts: Deliberately underspecified (e.g., “search for X,” “tell me about Y”).
2.2 Task Evolution Across Turns
The interaction progressed through:
- Identification of recent embassy-level remarks (finding none beyond a known date).
- Elevation of presidential announcements as primary sources.
- Collection of cabinet-level framing (USTR Davos keynote).
- Differentiation between primary government statements and secondary media reporting.
- Assembly of a compressed packet with executive summary, background, quotes, and use guidance.
At each step, ambiguity was preserved where evidence was absent, and claims were gated by attribution.
3. Outputs Produced
3.1 Executive Summary (Excerpt)
- No new standalone trade remarks by the Ambassador during the review period.
- A presidential announcement established headline terms of the trade deal.
- USTR remarks provided policy framing without India-specific commitments.
3.2 Structured Packet Components
- Background: Scope and limits of the agreement.
- Primary Quotes: Verbatim, attributable government statements.
- Secondary Reporting: Media context, clearly labeled.
- Use Guidance: How and where content may be reused safely.
The artifact was immediately usable for briefing, scripting, or internal reference without further normalization.
4. Comparative Analysis: Enterprise AI vs ACP
4.1 What Enterprise AI Can Replicate
- Multi-source search and summarization.
- Cross-turn synthesis within a session.
- Formatting into memos or briefs.
4.2 Failure Modes Observed in Enterprise AI
- Provenance drift: Media paraphrases elevated to primary claims.
- Over-compression: Loss of qualifiers (“reported,” “not India-specific”).
- Hallucinated completeness: Implied existence of remarks that were not found.
- Responsibility laundering: Outputs read as authoritative despite uncertainty.
4.3 ACP Differentiators
- Failure-as-output: Explicitly stating when no new remarks exist.
- Source stratification: Primary vs secondary vs inference.
- Authority discipline: Refusal to speak as an institution.
- Inspection survivability: Each claim traceable to a source or labeled as context.
5. Why Prompt Engineering Was Insufficient
Although no explicit governance prompt was provided, the interaction succeeded due to expert user intervention and conversational correction. This does not scale. Prompt-only approaches cannot:
- Enforce schemas once outputs leave the session.
- Prevent reuse without caveats.
- Guarantee refusal under pressure.
Prompts express intent; they do not enforce compliance.
6. Enforcing Governance at Scale
The case indicates that reliable deployment requires layers beyond prompts:
- Schemas: Mandatory sections (Primary Sources, Secondary Reporting, What Was Not Found).
- Tool Constraints: Source whitelists, citation requirements.
- Logging: URLs, timestamps, snapshots.
- Human Interrupts: Approval for quote tables and summaries.
7. Implications
For USG and regulated enterprises, the distinction is not model quality but governance posture. Systems that optimize for helpfulness risk institutional harm. Systems that optimize for restraint preserve human authority and accountability.
8. Conclusion
This case study demonstrates that ACP is not a stylistic variant of enterprise AI but a different operating discipline. Enterprise AI can approximate outputs; ACP preserves responsibility. The difference becomes visible only when tasks are underspecified, time-constrained, and politically sensitive—precisely the conditions under which institutions most need restraint.
Appendix: Reproducibility Notes
- All sources were public and unclassified.
- Quotes were verbatim and attributable at time of compilation.
- Absence of evidence was treated as a reportable outcome.
Appendix B: Methods — Comparative Execution Pathways
This appendix documents the methodological steps used in the case study and contrasts how the same task would typically be executed under (1) a standard enterprise AI workflow and (2) an ACP-governed workflow. The purpose is not to evaluate model capability, but to surface governance-relevant differences in process, failure handling, and artifact quality.
B.1 Task Definition and Constraints
Task: Assemble a rapid-turn, unclassified briefing packet on recent U.S.–India trade developments suitable for U.S. government personnel.
Key characteristics of the task:
- Underspecified initial prompts (no predefined format, scope, or citation rules).
- Time-sensitive and politically salient subject matter.
- Requirement for attribution-safe reuse.
- Mixed source environment (official government statements, media reporting, absence of new remarks).
These characteristics are representative of real-world institutional research requests.
B.2 Typical Enterprise AI Execution Path
In a standard enterprise AI setting, the task would proceed as follows:
- Broad retrieval: The system performs keyword-based search across news, web, and selected government sites.
- Semantic aggregation: Retrieved content is summarized into a unified narrative optimized for coherence and readability.
- Implicit normalization: Media reporting, paraphrases, and official statements are blended unless explicitly separated by the user.
- Confidence optimization: The system fills gaps with inferred continuity (e.g., implying recent remarks where none exist).
- Single-output delivery: A clean, fluent summary is produced, typically without an explicit section for non-findings.
Observed failure modes:
- Primary and secondary sources are not consistently distinguished.
- Absence of evidence is silently converted into assumed presence.
- Qualifiers are dropped during compression.
- The final artifact reads as authoritative despite unresolved uncertainty.
These failures are structural, not model-specific, and persist even with high-quality prompts.
B.3 ACP-Governed Execution Path (This Case)
Under the ACP posture demonstrated in this interaction, the task followed a different method:
- Progressive scoping: Each turn narrowed scope based on what was found and what was explicitly not found.
- Source stratification: Content was continuously categorized as:
- Primary U.S. government sources,
- Secondary media reporting,
- Contextual inference.
- Failure-as-output: The absence of recent embassy-level remarks was treated as a valid and reportable result.
- Quote discipline: Verbatim quotations were isolated, attributed, and separated from narrative synthesis.
- Packaging for inspection: The final packet included executive summary, background, quotes, media context, and use guidance.
At no point was uncertainty smoothed for readability alone.
B.4 Role of Prompting in This Case
Notably, the user did not provide a formal governance prompt at the outset. Governance emerged through:
- Expert user correction across turns,
- Explicit requests for attribution and compression discipline,
- Acceptance of partial and negative findings.
This demonstrates that expert users can temporarily act as a governance layer, but this reliance is fragile and non-scalable.
B.5 Implications for System Design
The case indicates that reliable institutional use requires mechanisms beyond prompt engineering:
- Schemas: Mandatory sections such as “Primary Sources,” “Secondary Reporting,” and “What Was Not Found.”
- Tool constraints: Source whitelists and citation requirements enforced at the system level.
- Auditability: Logged URLs, timestamps, and retrieval context.
- Human interruption: Clear authority to halt or revise outputs before dissemination.
Prompt engineering can express intent, but enforcement requires system-level design.
B.6 Replicability
This method can be replicated by:
- Running the same task under a standard enterprise AI system without governance constraints.
- Comparing whether the system:
- Invents recent remarks,
- Collapses media and official sources,
- Omits explicit non-findings.
The divergence in outputs provides an empirical basis for evaluating governance posture rather than model fluency.
Member discussion: