Loud or Silent? A Reusable Framework for Per-Modality Failure Analysis in Multimodal Clinical AI

📅 2026-08-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Multimodal clinical AI systems often suffer performance degradation in deployment due to missing modalities, yet it remains challenging to identify which modality fails and whether its failure is detectable (“loud”) or undetectable (“silent”). This work proposes a model-agnostic, fine-grained failure analysis framework that relies solely on observable signals during deployment. By leveraging modality embeddings, mask-aware probes, and labels, the method enables per-sample, per-modality failure attribution and outputs both a modality complementarity matrix and loud/silent failure profiles. In simulated data, the framework accurately recovers predefined modality dominance relationships; in the real-world MIMIC-IV cohort, it reveals that missing echocardiography nearly doubles error rates and, for the first time, quantifies the proportions of monitorable versus unmonitorable failures across modalities.
📝 Abstract
Multimodal clinical models are usually judged on accuracy with every modality present, but deployment removes modalities; an echocardiogram is often unavailable where an ECG is routine. Two questions then matter beyond the size of the accuracy loss: which modality was responsible, and whether the model fails loudly or silently once that modality is dropped. The distinction is per-example and modality-level, and is separate from post-hoc feature attribution (e.g. SHAP). Models are replaced often; the evaluation that answers these questions is reused. We present a model-agnostic modality-failure framework: given N modality embeddings, any mask-aware probe, and labels, it returns a per-example failure taxonomy, a per-modality complementarity matrix that attributes error to modalities, and a loud-vs-silent dropout profile separating monitorable failures from those that pass unflagged far from the decision boundary, using only deployment-observable signals. We release it as a small, unit-tested harness and validate it against planted ground truth. Across seeds it recovers that planted modality dominance and complementary subset, reports per-modality loud-vs-silent rates, and scales to a three-modality complementarity matrix; because the planted structure is known by construction, this validates recovery of per-example attribution rather than clinical performance. We then instantiate the framework on frozen EchoJEPA and HuBERT-ECG embeddings for LVEF and the EF <= 40% HFrEF gate over a paired MIMIC-IV cohort, where on the held-out test split (n = 245) dropping echo nearly doubles error. The narrow echo-to-ECG overlap that bounds cohort size is itself a deployment finding for cardiac foundation models. All of our work can be found at https://github.com/criticaldata/PRIMED-AI.
Problem

Research questions and friction points this paper is trying to address.

multimodal clinical AI
modality failure
loud vs silent failure
failure attribution
deployment robustness
Innovation

Methods, ideas, or system contributions that make the work stand out.

modality failure analysis
loud vs silent failure
model-agnostic framework
complementarity matrix
multimodal clinical AI
💼 Related Jobs
No related jobs found.
Q
Quang Bui
MIT Critical Data, Massachusetts Institute of Technology, Cambridge, Massachusetts, United States
S
Shlok Jaiswal
Neuqua Valley High School, Naperville, Illinois, United States
S
Samuel Paik-Heintz
North Hollywood High School, Los Angeles, California, United States
K
Kevin Zhou
Hopewell Valley Central High School, Pennington, New Jersey, United States
K
Kaushik Madapati
MIT Critical Data, Massachusetts Institute of Technology, Cambridge, Massachusetts, United States; Department of Electrical Engineering and Computer Sciences, University of California, Berkeley, Berkeley, California, United States
K
Krittaphas Chaisutyakorn
Siriraj Informatics and Data Innovation Center, Faculty of Medicine, Siriraj Hospital, Bangkok, Thailand; Harvard T.H. Chan School of Public Health, Harvard University, Boston, Massachusetts, United States
N
Noah Dane Hebdon
Quantum Innovation Centre (Q.InC), Agency for Science, Technology and Research, Singapore; School of Advanced International Studies, Johns Hopkins University, Washington, District of Columbia, United States
D
Dimitrios Proios
Department of Radiology and Medical Informatics, University of Geneva, Geneva, Switzerland
Sebastián Andrés Cajas Ordóñez
Sebastián Andrés Cajas Ordóñez
Harvard University
mhealthdeep learningcomputer visionaerospace
K
Kacper Dobek
MIT Critical Data, Massachusetts Institute of Technology, Cambridge, Massachusetts, United States; Institute of Computing Science, Poznan University of Technology, Poznan, Poland
Boya Zhang
Boya Zhang
Lawrence Livermore National Laboratory
Design of ExperimentsGaussian processesActive learning
A
Aly Dhedhi
American Heritage School, Plantation, Florida, United States
A
Ahram Han
MIT Critical Data, Massachusetts Institute of Technology, Cambridge, Massachusetts, United States; Institute for Medical Engineering and Science, Massachusetts Institute of Technology, Cambridge, Massachusetts, United States; Department of Surgery, Seoul National University Hospital, Seoul, South Korea
K
Kushul Reddy Palakala
University of North Florida, Jacksonville, Florida, United States
R
Rahul Gorijavolu
MIT Critical Data, Massachusetts Institute of Technology, Cambridge, Massachusetts, United States; School of Medicine, Johns Hopkins University, Baltimore, Maryland, United States; Department of Biomedical Engineering, Johns Hopkins University, Baltimore, Maryland, United States; Artificial Intelligence for Responsible, Generalizable, and Open Surgical (ARGOS) Research Group, Baltimore, Maryland, United States
J
Jacques Kpodonu
Division of Cardiac Surgery, Beth Israel Deaconess Medical Center, Harvard Medical School, Boston, Massachusetts, United States
Leo Anthony Celi
Leo Anthony Celi
Massachusetts Institute of Technology