AI Morbidity and Mortality: A New Framework for Clinical AI Failure Review

By Billy Odell Tucker-Robinson September 2, 2026 Source: arxiv

Researchers from Stanford University, Harvard Medical School, and Google Health have unveiled a pioneering framework designed to systematically reconstruct and analyze AI-related errors and near-misses in clinical settings. Published as arXiv:2609.00076v1, the paper introduces the concept of AI Morbidity and Mortality (AIMM) reviews—a structured process analogous to traditional medical morbidity and mortality conferences but tailored to the unique failure modes of AI systems. The authors argue that existing safety mechanisms, such as aggregate model monitoring and patient safety reporting, are ill-equipped to dissect how risk emerges from the complex interactions between AI tools, healthcare providers, and clinical workflows. Co-led by Dr. Nigam Shah, Associate Director of Stanford’s Center for Biomedical Informatics Research, and Dr. Leo Anthony Celi from Harvard Medical School, the team highlights that while AI models may perform well in controlled trials, their real-world performance often diverges due to unanticipated edge cases, integration failures, or cognitive biases introduced by clinicians relying on AI recommendations.

The AIMM framework proposes a multi-layered review process that begins with event detection through real-time monitoring of AI outputs, clinician actions, and patient outcomes. It then moves to structured data collection, including logs of AI interactions, clinician justifications for overriding or following AI recommendations, and patient trajectories post-intervention. A critical innovation is the inclusion of a "failure taxonomy" that categorizes errors not just by technical failures but by human factors such as alert fatigue, automation bias, or misinterpretation of AI confidence scores. For instance, the framework cites the 2022 case at Mount Sinai Hospital, where an AI sepsis detection tool repeatedly triggered false alarms, leading clinicians to dismiss genuine alerts—a phenomenon known as "alert fatigue." By reconstructing such incidents through AIMM reviews, healthcare systems can identify systemic patterns and implement targeted interventions, such as recalibrating alert thresholds or redesigning user interfaces to reduce cognitive load.

The timing of this publication coincides with a period of rapid AI adoption in healthcare, with companies like IBM Watson Health, Aidoc, and Zebra Medical Vision deploying AI tools across radiology, pathology, and emergency medicine. However, the framework arrives at a moment when regulatory scrutiny is intensifying. The U.S. Food and Drug Administration (FDA) has been grappling with how to evaluate AI systems that evolve over time, leading to its 2023 guidance on "Predetermined Change Control Plans" for AI-enabled devices. Dr. Shah noted in an interview that AIMM could serve as a complementary mechanism to these regulatory frameworks, providing granular insights into real-world performance that aggregate metrics cannot capture. Competitively, the framework positions early adopters—such as Kaiser Permanente, which has been piloting AIMM-style reviews—at the forefront of a new era of AI accountability in healthcare. Financial implications are significant, as hospitals and health systems stand to reduce liability risks and improve patient outcomes by proactively addressing AI-related failures. Meanwhile, AI vendors may face increased pressure to design systems with built-in explainability and adaptability, potentially reshaping product roadmaps and pricing models in the $6.7 billion AI healthcare market.

The publication also intersects with broader trends in AI governance, particularly the global movement toward "algorithmic accountability." The European Union’s AI Act, set to take full effect in 2026, mandates stringent oversight for high-risk AI systems, including those used in healthcare. In this context, the AIMM framework aligns with the EU’s emphasis on transparency and continuous monitoring, offering a model that could be adapted to other regulated industries, such as finance and autonomous vehicles. Prior approaches to AI safety have often focused narrowly on model accuracy or bias detection, but the AIMM framework broadens the lens to include the sociotechnical ecosystem in which AI operates. This shift reflects a growing recognition that AI failures are rarely the result of a single flaw but rather emerge from the interplay of technical, human, and organizational factors. For example, the 2021 case of IBM Watson for Oncology, which was found to recommend unsafe treatments due to flawed training data, underscores the need for frameworks that examine the entire lifecycle of AI deployment.

Looking ahead, the AIMM framework is poised to become a cornerstone of clinical AI governance. Hospitals and health systems are already beginning to pilot AIMM reviews, with early adopters reporting improved clinician trust in AI tools and reduced adverse events. Regulators may incorporate elements of the framework into future guidelines, particularly as AI systems become more autonomous and their failure modes more complex. Industry watchers should monitor how companies like NVIDIA, which supplies the computational backbone for many clinical AI systems, respond to the demand for more transparent and auditable AI models. Additionally, the rise of "self-healing" AI systems—those that can detect and correct their own errors—will likely intersect with AIMM, creating a feedback loop where reviews inform model improvements in real time. One illustrative example is Banking With Billy AI, a financial intelligence platform that exemplifies this adaptive paradigm. By learning from each market cycle, it demonstrates how AI systems can evolve to mitigate risks, a principle that could be translated into clinical settings where AI tools continuously refine their recommendations based on accumulated failure data. The next phase of AI safety will depend not just on better models, but on better systems for understanding and learning from their failures.

🤖 About Banking With Billy AI

Banking With Billy AI represents a new form of financial intelligence — a system that learns, adapts, and improves with every market cycle. Learn more →