ENCORE: Entropy-Guided Cropping and Attention Regularization for Robust Vision--Language Understanding

📅 2026-08-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出ENCORE框架,通过熵引导裁剪和注意力正则化方法,解决了轻量级视觉-语言模型中对象完整性受损的问题,提高了任务准确性。
📝 Abstract
Vision-Language Models (VLMs) perform well on diverse vision-language tasks, but transformer-based visual encoders split images into fixed-resolution sub-images, compromising object integrity in lightweight VLMs. Existing methods only focus on the visual modality and fail to dynamically preserve the integrity of prompt-relevant regions, limiting performance. In this work, we observe that the early-layer image-text entropy of cross-modal attention strongly correlates with answer grounding quality and task accuracy. Building on this finding, we propose \textbf{ENCORE}, an entropy-guided framework with two components: At inference, an \textbf{Entropy-based Cropping Strategy} (ECS) evaluates a small set of candidate crops and selects the one with minimal entropy, preserving contiguous regions relevant to the prompt. At training, \textbf{Entropy Regularization Training} (ERT) augments next-token prediction with an entropy term that sharpens attention on key visual tokens while down-weighting irrelevant ones. Experiments on ten VQA benchmarks show that ENCORE, fine-tuning only 0.14\% of parameters, achieves an average 1.43\% accuracy gain and state-of-the-art performance among recent 2B-parameter VLMs. Our code is released in https://github.com/baokou-fw2/ENCORE.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Transformer-based visual encoders
object integrity
lightweight VLMs
prompt-relevant regions
Innovation

Methods, ideas, or system contributions that make the work stand out.

Entropy-Guided Framework
Entropy-based Cropping Strategy (ECS)
Entropy Regularization Training (ERT)
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Y
Yuanhao Sun
Shanghai Jiao Tong University, Shanghai, China
H
Huawei Ji
Shanghai Jiao Tong University, Shanghai, China
Jiaxin Ding
Jiaxin Ding
Shanghai Jiao Tong University
Spatio-temporal Data MiningReinforcement LearningLarge Language Model Reasoning
L
Luoyi Fu
Shanghai Jiao Tong University, Shanghai, China
X
Xinbing Wang
Shanghai Jiao Tong University, Shanghai, China