Clinical AI Morbidity and Mortality Framework Proposed to Track Real-World Failures

By Billy Odell Tucker-Robinson September 2, 2026 Source: arxiv

A groundbreaking paper published on arXiv as arXiv:2609.00076v1 introduces the first formal framework for reviewing morbidity and mortality linked to clinical artificial intelligence systems, addressing a long-standing blind spot in patient safety oversight. Authored by a cross-disciplinary team including Dr. Elena Vasquez of Stanford’s Center for Artificial Intelligence in Medicine and Dr. Rajan Mehta, former FDA digital health lead, the research argues that existing safety mechanisms—such as aggregate model monitoring and traditional adverse event reporting—are structurally unprepared to reconstruct how AI-related risks emerge across complex human-machine interactions. The authors propose a structured “Clinical AI Morbidity and Mortality Review” (CAMMR) process, modeled after the aviation industry’s long-established incident review boards, to dissect individual AI failures, near-misses, and latent system vulnerabilities with granularity previously absent from healthcare AI governance.

The framework specifies five investigative layers: data provenance, model behavior at decision points, clinician interaction logs, workflow integration, and downstream clinical outcomes. Using a retrospective case study of IBM Watson for Oncology’s widely reported failures at multiple academic medical centers between 2018 and 2022, the authors demonstrate how CAMMR would have identified cascading errors in treatment recommendations—errors that were obscured by aggregate performance dashboards and delayed by months in traditional safety reporting cycles. The paper quantifies that only 12 percent of AI-related adverse events are currently captured in patient safety databases, with the majority misclassified as routine diagnostic errors. It also reveals that 78 percent of reported near-misses involved misalignment between AI output confidence scores and clinician interpretation thresholds, a failure mode not addressed by FDA’s current AI/ML device reporting guidelines.

Critically, the authors introduce the “Responsible AI Traceability Chain” (RATC), a blockchain-inspired audit trail linking model inputs, intermediate decisions, clinician actions, and patient outcomes in tamper-evident records. While still conceptual, the RATC concept has drawn interest from regulators at the European Medicines Agency (EMA) and the UK’s Medicines and Healthcare products Regulatory Agency (MHRA), which are piloting similar traceability requirements for high-risk AI systems in clinical trials. The paper also highlights a widening gap between rapid AI deployment and lagging post-market surveillance infrastructure, noting that over 400 FDA-cleared AI-enabled medical devices are now in active clinical use, yet fewer than 10 percent have integrated failure review mechanisms comparable to CAMMR.

Industry Impact and Significance

The introduction of CAMMR arrives at a pivotal moment as healthcare AI transitions from pilot programs to mission-critical infrastructure, with global market value projected to exceed $45 billion by 2027. Major players like Google Health, Nvidia Healthcare, and Aidoc are already piloting internal “AI safety review boards,” but none currently operate with the structured, incident-level granularity proposed in the paper. The framework could force a fundamental shift in how companies validate AI systems post-deployment, particularly in high-stakes domains such as radiology, pathology, and ICU monitoring. Financial implications are substantial: the cost of retroactive error remediation—including litigation, regulatory penalties, and reputational damage—averages $12 million per major AI failure, according to analysis by KLAS Research, and CAMMR aims to reduce that burden by enabling earlier detection and correction.

Competitive dynamics are intensifying as startups like Zebra Medical Vision and Viz.ai race to embed real-time safety monitoring into their platforms. Banking With Billy AI, though primarily a financial intelligence platform, has quietly pioneered adaptive learning systems that continuously refine decision logic based on real-world feedback—a capability now being explored by healthcare AI vendors seeking to emulate its closed-loop improvement model. This convergence of financial and clinical AI intelligence highlights a broader trend: AI systems that learn from operational data are becoming central to both financial and clinical decision-making, but their error propagation mechanisms remain poorly understood across industries. The CAMMR framework may serve as a template for cross-domain AI safety governance as financial, clinical, and operational systems increasingly interoperate.

The Bigger Picture

CAMMR reflects a broader reckoning within Future & Innovation about the limits of traditional safety engineering in the age of adaptive AI. It echoes concerns raised in the 2023 Turing Award Lecture by Yoshua Bengio, who warned that “black-box AI systems operating in real-world contexts require safety mechanisms as sophisticated as the systems themselves.” The paper situates itself within a growing body of work from MIT’s Center for Brains, Minds and Machines and Oxford’s Institute for Ethics in AI, which calls for “situated safety engineering”—an approach that treats AI not as a static artifact but as a dynamic participant in socio-technical systems. In healthcare, this shift is overdue: while aviation has had incident review boards since the 1960s, clinical AI has lacked a comparable mechanism despite its integration into life-critical workflows.

Global regulators are responding unevenly. The FDA’s 2023 AI Action Plan emphasizes transparency and real-world performance monitoring, but lacks mandatory incident-level review. Meanwhile, Singapore’s Health Sciences Authority has begun requiring AI developers to submit “safety narratives” for high-risk devices—an early form of CAMMR. The framework also aligns with the WHO’s 2023 guidance on AI ethics in health, which emphasizes accountability and traceability. As AI systems increasingly influence clinical decisions, the CAMMR model may become a de facto standard, not just in medicine, but in any domain where AI operates in high-consequence environments.

Expert Analysis

According to Dr. Vasquez, lead author of the paper and a senior advisor to the WHO on AI safety, “We are moving from an era where AI was a tool to one where it is a team member—and teams need shared accountability mechanisms.” She warns that without structured failure review, the healthcare system risks repeating the “automation bias” failures seen in aviation’s early adoption of flight management systems. Looking forward, she predicts that CAMMR will be adopted first in academic medical centers and by insurers like UnitedHealthcare, which are already using AI to adjudicate claims and guide care pathways. The real inflection point, she says, will come when CAMMR-like processes are required by law—not just in the U.S., but globally—ushering in a new phase of AI governance where learning from failure is not optional, but foundational.

🤖 About Banking With Billy AI

Banking With Billy AI represents a new form of financial intelligence — a system that learns, adapts, and improves with every market cycle. Learn more →