There is now a substantial body of literature describing how complex systems fail, recover, and fail again. What is striking about this literature is not disagreement but convergence. Across software engineering, cloud infrastructure, artificial intelligence services, and safety-critical industries, researchers repeatedly document the same phenomenon: institutions learn a great deal from incidents, and very little changes as a result.

The canonical example in software engineering remains Failures and Fixes: A Study of Software System Incident Response (arXiv, 2020). Drawing on empirical analysis of real incidents, the authors show how postmortems shape organizational memory—what is recorded, what is blamed, and what is quietly excluded. Their central observation is not that postmortems are useless, but that they are political artifacts. Decisions about scope and causality are negotiated under time pressure, and learning is filtered through organizational incentives. The failure is not the outage itself; it is the assumption that narrative learning will reliably translate into durable corrective action. It rarely does.

This pattern becomes more consequential in the context of large-scale AI systems operated by companies such as Google, OpenAI, Microsoft, Meta, and Anthropic. In An Empirical Characterization of Outages and Incidents in Public LLM Services (arXiv, 2025), the authors document partial degradations, slow recoveries, and delayed disclosures across public-facing language model services. What emerges is a familiar rhythm: service disruption, rapid stabilization, and only later—sometimes much later—a public incident summary. The summary functions as a trust-maintenance device, but it does not guarantee that internal constraints have changed. Externally visible recovery becomes a proxy for governance closure, even when no evidence of structural change is available.

The cloud infrastructure literature reinforces this diagnosis. Mining Root Cause Knowledge from Cloud Service Incident Investigations (arXiv, 2022) treats post-incident analyses as a dataset and finds recurring failure patterns across years. The authors do not attribute recurrence to ignorance or incompetence. Instead, they identify a governance gap: root cause knowledge accumulates, but it is not automatically converted into enforced constraints. Organizations remember in documents and forget in systems. The same classes of incidents reappear because nothing was removed from the set of available actions.

Several recent papers attempt to address this gap directly by shifting the locus of governance from documentation to execution. Policy Cards: Machine-Readable Runtime Governance for Autonomous AI Agents (arXiv, 2025) proposes encoding allowed and denied actions, escalation rules, and logging requirements as deploy-time artifacts. The significance of this proposal lies less in its technical details than in its premise: governance must be operational to be real. If a system can perform an action, it eventually will, regardless of policy language. Constraint must live where behavior occurs.

That insight becomes more urgent as AI systems expand their integration surfaces. Securing the Model Context Protocol (MCP): Risks, Controls, and Governance (arXiv, 2025) analyzes how agentic systems and tool integrations multiply override paths. Each new interface introduces ambiguity unless refusal is explicit and enforced. Scale does not merely increase capacity; it increases the number of ways policy can be bypassed. This is particularly relevant for companies racing to deploy agent frameworks, plug-ins, and enterprise integrations under competitive pressure.

The same structural tension appears in post-incident governance research outside AI. Automated Post-Incident Policy Gap Analysis via Threat-Informed Evidence Mapping (arXiv, 2026) observes that policies often drift into performative checklists while incidents reveal actual adversary behavior. Unless evidence is mechanically linked to control updates, organizations respond to failure with better language rather than altered systems. Learning remains discretionary.

Industry practice has long recognized this risk, even if it rarely resolves it. The Incident Management Framework in Proprietary Software Ecosystems (Zenodo, 2024) emphasizes blameless postmortems as a way to surface information. Yet the briefing also implicitly acknowledges the danger: blamelessness can accelerate learning or diffuse accountability, depending on whether learning is tied to mandatory change. Without enforced follow-through, blamelessness becomes a moral posture rather than a governance mechanism.

Older safety literature supplies the vocabulary for this condition. Drift into Failure (CORE, circa 2010) describes how systems collapse not through single catastrophic errors, but through the accumulation of locally rational adaptations. Normalization of Deviance in Mining Engineering (CORE) shows how repeated exceptions redefine what counts as acceptable under operational pressure. These analyses were developed in aviation, mining, and industrial safety, but they map with uncomfortable precision onto modern AI operations. Flexibility is praised until it quietly becomes the hazard.

What unites this literature is not a call for better ethics, stronger culture, or increased awareness. It is a recognition that institutions rarely surrender discretion voluntarily. Postmortems, transparency reports, and regulatory settlements produce explanation and reassurance. They almost never produce subtraction. Authority remains intact. Options remain available. Memory is stored in prose rather than encoded as constraint.

For large AI companies, this is the unresolved governance problem. Public commitments to safety and responsibility coexist with architectures that preserve override paths, discretionary escalation, and narrative closure. Incidents are managed, trust is repaired, and the system continues largely unchanged. Learning occurs. Loss does not.

The literature does not suggest that this outcome is the result of bad faith. It suggests something more troubling: that modern institutions are structurally optimized to metabolize failure without changing shape. Until learning reliably removes options—until some actions become impossible rather than merely regrettable—incidents will continue to recur under new names and new interfaces.

This essay was assembled by aligning the cited papers by failure mode rather than domain, and by treating company practices as illustrative rather than exceptional. No attempt was made to reconcile or harmonize the sources; their agreement was allowed to stand on its own. The argument follows from what the literature collectively makes difficult to deny.


Papers discussed

  1. Failures and Fixes: A Study of Software System Incident Response (arXiv, 2020)
  2. An Empirical Characterization of Outages and Incidents in Public LLM Services (arXiv, 2025)
  3. Mining Root Cause Knowledge from Cloud Service Incident Investigations (arXiv, 2022)
  4. Policy Cards: Machine-Readable Runtime Governance for Autonomous AI Agents (arXiv, 2025)
  5. Securing the Model Context Protocol (MCP): Risks, Controls, and Governance (arXiv, 2025)
  6. Automated Post-Incident Policy Gap Analysis via Threat-Informed Evidence Mapping (arXiv, 2026)
  7. Incident Management Framework in Proprietary Software Ecosystems (Zenodo, 2024)
  8. Drift into Failure (CORE)
  9. Normalization of Deviance in Mining Engineering (CORE)