SeVeR: Selective Visual Exposure and Retrieval for 3D Medical Image Question Answering

📅 2026-08-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究解决了3D医学图像问答中的视觉冗余问题,通过SeVeR框架选择性地压缩和检索多模态信息,减少了不必要的视觉暴露,提高了性能。
📝 Abstract
Volumetric medical VQA requires reasoning over long and redundant 3D visual token sequences, especially in multi-sequence MRI where complementary modalities provide diverse diagnostic cues but expose the decoder to many repeated anatomical regions. To investigate reasoning under multi-sequence visual redundancy, we first introduce BreMRIs-VQA, a clinically curated breast MRI benchmark with 1.19M QA pairs from 71.0K sequences and 12.9K patients, covering both free-text and multiple-choice questions. We further propose SeVeR, a selective visual exposure framework that compresses dense volumes into modality-wise prototypes and retrieves complementary multi-level evidence with change-aware gated attention during decoding, trained with a marginal-utility self-consistency objective that suppresses unhelpful retrieval. Experiments on BreMRIs-VQA and public benchmarks show that SeVeR improves both discriminative and generative performance while exposing substantially fewer visual tokens.
Problem

Research questions and friction points this paper is trying to address.

3D Medical Image Question Answering
visual redundancy
multi-sequence MRI
Innovation

Methods, ideas, or system contributions that make the work stand out.

Selective Visual Exposure
Change-aware Gated Attention
Marginal-utility Self-consistency
Multi-sequence MRI
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yaojun Hu
DAMO Academy, Alibaba Group; College of Computer Science and Technology, Zhejiang University; State Key Laboratory of Transvascular Implantation Devices and TIDRI, Zhejiang University; The Second Affiliated Hospital and Liangzhu Laboratory, Zhejiang University School of Medicine; Zhejiang Key Laboratory of Medical Imaging Artificial Intelligence
D
Danyang Tu
DAMO Academy, Alibaba Group; Hupan Lab; College of Computer Science and Technology, Zhejiang University
Y
Yang Liu
DAMO Academy, Alibaba Group
Jiajin Zhang
Jiajin Zhang
Alibaba, DAMO Academy
Medical Image AnalysisComputer VisionMedical Imaging
Wei Fang
Wei Fang
Alibaba Damo Academy, Zhejiang University; Tsinghua University
CT reconstructionMedical Image AnalysisDeep learning
Zhiqiang Liu
Zhiqiang Liu
zhejiang university
C
Chunlai Dong
DAMO Academy, Alibaba Group; College of Computer Science and Technology, Zhejiang University; State Key Laboratory of Transvascular Implantation Devices and TIDRI, Zhejiang University; The Second Affiliated Hospital and Liangzhu Laboratory, Zhejiang University School of Medicine; Zhejiang Key Laboratory of Medical Imaging Artificial Intelligence
Y
Yingda Xia
DAMO Academy, Alibaba Group
H
Haochao Ying
School of Public Health, Zhejiang University; State Key Laboratory of Transvascular Implantation Devices and TIDRI, Zhejiang University; The Second Affiliated Hospital and Liangzhu Laboratory, Zhejiang University School of Medicine; Zhejiang Key Laboratory of Medical Imaging Artificial Intelligence
J
Jian Wu
DAMO Academy, Alibaba Group; College of Computer Science and Technology, Zhejiang University; State Key Laboratory of Transvascular Implantation Devices and TIDRI, Zhejiang University; The Second Affiliated Hospital and Liangzhu Laboratory, Zhejiang University School of Medicine; Zhejiang Key Laboratory of Medical Imaging Artificial Intelligence
Ling Zhang
Ling Zhang
Alibaba DAMO Academy USA
Medical Image AnalysisMedical Image ComputingMachine LearningImage Processing