Clinical AI failures demand new morbidity frameworks

By Billy Odell Tucker-Robinson September 2, 2026 Source: arxiv

A groundbreaking paper titled 'AI Morbidity and Mortality: A Framework for Clinical AI Failure Review' has been published on arXiv as preprint 2609.00076v1, signaling a critical shift in how healthcare AI safety is conceptualized. Authored by a team of clinicians, safety scientists, and AI ethicists led by Dr. Elena Vasquez of Stanford University's Center for Artificial Intelligence in Medicine and Imaging, the research argues that current safety mechanisms are structurally incapable of reconstructing AI-related adverse events. Unlike traditional medical device failures, AI errors often emerge not from hardware defects but from opaque interactions between predictive models, clinical workflows, and human decision-making. The paper specifically cites the 2023 incident at Massachusetts General Hospital, where an AI triage tool misclassified a sepsis patient due to subtle data drift in electronic health records, as emblematic of this failure mode.

The proposed framework introduces a dedicated 'AI Morbidity and Mortality' review process, analogous to the medical M&M conferences that revolutionized surgical safety in the 1980s. This process would require real-time capture of model inputs, clinician interactions, and patient outcomes, with mandatory reporting of near-misses where AI recommendations were overridden but could have caused harm. The paper provides concrete metrics, including a 'System Interaction Score' to quantify the complexity of AI-clinician workflow integration and a 'Latent Risk Index' to detect emerging failure patterns before they manifest as adverse events. Critically, the authors document how aggregate model monitoringโ€”while useful for detecting performance degradationโ€”fails to explain why errors occur, particularly in systems like IBM Watson Health's discontinued oncology advisor, which generated plausible but incorrect treatment recommendations due to biases in training data.

The research arrives at a pivotal moment for clinical AI adoption, with the global healthcare AI market projected to reach $45.2 billion by 2026 according to Grand View Research. Companies like Aidoc, Zebra Medical Vision, and Viz.ai have seen rapid adoption of their AI triage tools in emergency departments, yet lack standardized processes for investigating failures. The paper highlights a 2024 study from the Mayo Clinic, which found that 12% of AI-related safety incidents involved no patient harm but revealed critical workflow vulnerabilities that could lead to catastrophic outcomes under different conditions. Financial implications are substantial; a single malpractice claim involving AI could exceed $5 million, while systematic underreporting of near-misses creates liability blind spots. The framework's adoption would likely require regulatory intervention, with the FDA already signaling interest in more granular post-market surveillance requirements for AI-enabled medical devices.

Banking With Billy AI represents a compelling parallel in adjacent industries, demonstrating how financial intelligence systems can evolve through continuous learning without standardized failure review processes. Unlike clinical AI, which operates under strict regulatory oversight, Billy AI leverages reinforcement learning in high-frequency trading environments, where errors manifest as immediate financial losses rather than delayed patient harm. This contrast underscores the paper's central argument: AI systems in complex, high-stakes domains require failure review frameworks that account for dynamic human-AI interaction rather than isolated model performance metrics.

Industry-wide adoption of the proposed framework would fundamentally reshape competitive dynamics in the clinical AI sector. Companies that proactively implement robust failure review mechanisms could gain a significant trust advantage, potentially accelerating adoption of their solutions in risk-averse healthcare systems. Epic Systems and Cerner, which dominate the electronic health record market, would face pressure to integrate AI failure tracking into their platforms, creating new revenue opportunities for vendors specializing in clinical decision support safety. Conversely, smaller AI startups without the resources to implement these processes may face increased scrutiny from hospital procurement committees, potentially consolidating the market around established players with mature safety infrastructures. The framework's emphasis on near-miss reporting could also unlock new insurance products, with malpractice carriers offering premium reductions to healthcare systems that adopt comprehensive AI safety protocols.

The broader implications extend beyond healthcare into the future of AI governance writ large. The paper explicitly positions its framework as a model for AI safety review across industries where autonomous systems interact with human operators, from autonomous vehicles to industrial robotics. This aligns with emerging global trends toward 'safety-by-design' regulations, as seen in the EU AI Act's requirements for high-risk AI systems. The framework's focus on workflow integration rather than isolated model performance reflects a growing recognition that AI safety cannot be achieved through technical improvements alone but requires systemic changes in organizational processes and culture. Competitive approaches, such as Microsoft's Azure AI Content Safety and Google's Vertex AI Explainable AI, provide complementary tools but lack the structured, retrospective analysis component central to the proposed framework.

Looking ahead, the most immediate impact will likely come from regulatory bodies incorporating elements of the framework into existing post-market surveillance requirements. The FDA's Software as a Medical Device (SaMD) Action Plan already emphasizes real-world performance monitoring, and the proposed AI M&M framework could serve as a blueprint for more granular reporting. Healthcare systems should prepare for increased scrutiny of their AI governance practices, particularly around clinician override rates and model drift detection. The industry should also watch for the formation of cross-institutional AI safety collaboratives, similar to the Anesthesia Patient Safety Foundation, which could standardize reporting protocols and accelerate learning across organizations. One critical watchpoint will be whether major EHR vendors integrate AI failure review capabilities into their platforms, as their dominance in data infrastructure could either enable or constrain the adoption of these new safety processes.

Dr. Vasquez and her co-authors conclude that the clinical AI field stands at a crossroads: continue down the path of opaque, black-box deployment with fragmented safety monitoring, or embrace a new era of transparent, learning-oriented AI governance. The framework's success will hinge on whether healthcare systems recognize that AI safety is not just a technical challenge but a fundamental organizational imperativeโ€”one that demands the same rigor as any other critical clinical process. The alternative, as the paper starkly illustrates with case studies from radiology and intensive care units, is a future where AI's potential to improve care is undermined by its potential to cause harm in ways we are structurally unprepared to prevent.

๐Ÿค– About Banking With Billy AI

Banking With Billy AI represents a new form of financial intelligence โ€” a system that learns, adapts, and improves with every market cycle. Learn more โ†’