BrainFocus: EEG-Guided ROI Selection for Efficient Vision-Language Models

📅 2026-09-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出BrainFocus框架,利用EEG信号和YOLO检测器选择相关图像区域,以提高视觉-语言模型在视觉问答任务中的效率和准确性。
📝 Abstract
Vision-language models (VLMs) achieve strong visual question answering (VQA) performance, but processing large cluttered images is computationally expensive when only a small region is relevant. Electroencephalography (EEG) signals, which capture human neural responses to visual stimuli, can provide a human-derived semantic cue about the region of interest (ROI). However, EEG-guided visual category decoding remains imperfect, making direct ROI routing unreliable. In this work, we propose BrainFocus, a reliable EEG-guided efficient VLM framework for VQA. An EEG classifier predicts a target category, and a YOLO detector localizes the matching ROI. The VLM receives the cropped ROI only when both predictions pass confidence thresholds; otherwise, it processes the full image. For evaluation, we build on EEG-ImageNet to construct a 40-class benchmark comprising generated cluttered images and real object-centric images, with target-ROI annotations and 600 English visual question-answer pairs. Across Qwen3.5-VL 2B, 4B, and 9B models, BrainFocus improves VQA accuracy by 4.14-9.87 percentage points (pp) on cluttered scenes while reducing input tokens and total tokens by 23.2%-39.4% and 23.2%-39.3%, and end-to-end floating-point operations (FLOPs) by 23.2%-39.5%. These results demonstrate that EEG can guide efficient VLM inference even when its semantic decoding is imperfect.
Problem

Research questions and friction points this paper is trying to address.

EEG
ROI
VLM
VQA
visual stimuli
Innovation

Methods, ideas, or system contributions that make the work stand out.

EEG-guided
Efficient VLM
ROI selection
Visual Question Answering
Confidence Threshold
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yihui Peng
Leiden Institute of Advanced Computer Science (LIACS), Leiden University, The Netherlands
G
Guorui Lu
Leiden Institute of Advanced Computer Science (LIACS), Leiden University, The Netherlands
Qinyu Chen
Qinyu Chen
Assistant Professor, Leiden University
Edge AIIC designNeuromorphic ComputingEvent-based visionAR/VR