6 Fingers, 1 Kidney: Natural Adversarial Medical Images Reveal Critical Weaknesses of Vision-Language Models

📅 2025-12-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work identifies a critical limitation in vision-language models (VLMs): severely degraded generalization under rare anatomical variations—stemming from strong prior biases toward “typical” anatomy. To address this, the authors introduce the novel concept of *natural adversarial anatomy* and construct AdversarialAnatomyBench, the first cross-modal, multi-anatomical-region benchmark for rare anatomical variants. Evaluating 22 state-of-the-art VLMs on fundamental medical perception tasks, they find that average accuracy plummets to 29% on atypical anatomy—down 45 percentage points from 74% on typical anatomy—with even the best-performing model suffering 41–51% relative performance loss. Standard mitigation strategies—including bias-aware prompting and test-time inference—yield negligible improvements. This study provides the first quantitative characterization and root-cause attribution of VLMs’ clinical robustness deficits, establishing a foundational evaluation paradigm and actionable directions for developing trustworthy medical AI.

Technology Category

Application Category

📝 Abstract
Vision-language models are increasingly integrated into clinical workflows. However, existing benchmarks primarily assess performance on common anatomical presentations and fail to capture the challenges posed by rare variants. To address this gap, we introduce AdversarialAnatomyBench, the first benchmark comprising naturally occurring rare anatomical variants across diverse imaging modalities and anatomical regions. We call such variants that violate learned priors about "typical" human anatomy natural adversarial anatomy. Benchmarking 22 state-of-the-art VLMs with AdversarialAnatomyBench yielded three key insights. First, when queried with basic medical perception tasks, mean accuracy dropped from 74% on typical to 29% on atypical anatomy. Even the best-performing models, GPT-5, Gemini 2.5 Pro, and Llama 4 Maverick, showed performance drops of 41-51%. Second, model errors closely mirrored expected anatomical biases. Third, neither model scaling nor interventions, including bias-aware prompting and test-time reasoning, resolved these issues. These findings highlight a critical and previously unquantified limitation in current VLM: their poor generalization to rare anatomical presentations. AdversarialAnatomyBench provides a foundation for systematically measuring and mitigating anatomical bias in multimodal medical AI systems.
Problem

Research questions and friction points this paper is trying to address.

Benchmark tests VLMs on rare anatomical variants
Reveals significant accuracy drop in atypical anatomy
Highlights poor generalization to rare medical presentations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Introduces AdversarialAnatomyBench for rare anatomical variants
Reveals VLMs' poor generalization to atypical anatomy
Shows scaling and prompting fail to resolve biases
🔎 Similar Papers
2024-07-02Conference on Empirical Methods in Natural Language ProcessingCitations: 1
Leon Mayer
Leon Mayer
PhD Student, German Cancer Research Center (DKFZ)
P
Piotr Kalinowski
German Cancer Research Center (DKFZ) Heidelberg, Division of Intelligent Medical Systems, Germany.
C
Caroline Ebersbach
German Cancer Research Center (DKFZ) Heidelberg, Division of Intelligent Medical Systems, Germany.
M
Marcel Knopp
German Cancer Research Center (DKFZ) Heidelberg, Division of Intelligent Medical Systems, Germany.
T
Tim Rädsch
German Cancer Research Center (DKFZ) Heidelberg, Division of Intelligent Medical Systems, Germany.
E
Evangelia Christodoulou
German Cancer Research Center (DKFZ) Heidelberg, Division of Intelligent Medical Systems, Germany.
A
Annika Reinke
German Cancer Research Center (DKFZ) Heidelberg, Division of Intelligent Medical Systems, Germany.
Fiona R. Kolbinger
Fiona R. Kolbinger
Purdue University
Surgical AISurgical Data ScienceMedical Image AnalysisSurgical Oncology
L
Lena Maier-Hein
German Cancer Research Center (DKFZ) Heidelberg, Division of Intelligent Medical Systems, Germany.