MedFG-VQA: Low-Frequency Memory and Graph Attention for Lightweight Medical VQA

📅 2026-08-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决医疗VQA中标注数据有限和计算需求高的问题,提出MedFG-VQA框架,利用低频记忆增强和图注意力机制实现高效视觉-文本对齐。
📝 Abstract
Medical Visual Question Answering (Med-VQA) holds significant promise for clinical decision support, yet faces challenges due to limited annotated data and the high computational demands of existing large vision-language models. We propose MedFG-VQA, a lightweight framework that leverages a memory bank to augment DCT-based low-frequency features and employs graph-enhanced cross-attention for effective visual-textual alignment. Specifically, our approach features two key components: Frequency-Memory Fusion (FMF), which enhances low-frequency features by retrieving from a learnable memory bank built on DCT decomposition, and Graph-Aware Cross-Attention (GACA), which aligns visual-textual features via cross-attention and refines them through graph-convolutional aggregation. To address data scarcity, we construct SynMed-VQA, a large-scale synthetic dataset comprising over 2 million question-answer pairs across 9 imaging modalities and 10 major organs, generated with GPT-4o. Extensive experiments on SynMed-VQA and three other standard biomedical VQA benchmarks demonstrate that MedFG-VQA achieves competitive or superior performance compared to much larger models while maintaining significantly lower computational costs, highlighting its efficiency and potential for clinical deployment.
Problem

Research questions and friction points this paper is trying to address.

Medical Visual Question Answering
annotated data
computational demands
Innovation

Methods, ideas, or system contributions that make the work stand out.

Frequency-Memory Fusion
Graph-Aware Cross-Attention
Lightweight Framework
Low-Frequency Features
Synthetic Dataset
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
H
Haowen Gu
School of Computer Science and Engineering, Nanjing University of Science and Technology; State Key Laboratory of Intelligent Manufacturing of Advanced Construction Machinery
G
Gensheng Pei
Department of Electrical and Computer Engineering, Sungkyunkwan University
Zeren Sun
Zeren Sun
Associate Professor, Nanjing University of Science and Technology
computer visiondeep learningfine-grained visual recognitionlearning from label noise
M
Mingwu Ren
School of Computer Science and Engineering, Nanjing University of Science and Technology; State Key Laboratory of Intelligent Manufacturing of Advanced Construction Machinery
X
Xiangbo Shu
School of Computer Science and Engineering, Nanjing University of Science and Technology
Y
Yazhou Yao
School of Computer Science and Engineering, Nanjing University of Science and Technology; State Key Laboratory of Intelligent Manufacturing of Advanced Construction Machinery
F
Fumin Shen
School of Computer Science and Engineering, University of Electronic Science and Technology of China