Learning to Look Again: Loss-Gap Supervision for Free-form Crop Routing in Vision-Language Models

📅 2026-08-22
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决视觉-语言模型在细节问题上的失败,提出GapSight框架,通过损失差距监督学习选择性重读图像局部区域的方法,提高模型性能。
📝 Abstract
Vision-language models (VLMs) fail many detail-centric questions for a concrete reason: the answer is visible in the image, yet lost after the image is compressed into a low-resolution global view. Allocating more visual tokens to every query improves some OCR and document cases, but it spends computation indiscriminately and can disturb tasks that rely on global context. We propose GapSight, a framework for learning visual re-reading: a VLM first takes a global glance, then selectively returns to a free-form region when the question calls for local evidence. The supervision comes from the target model's own failure signal. Offline, we compare answer loss or multiple-choice option margin under a global-only view and candidate crop-augmented views; crops that improve the target answer become model-specific review labels. A lightweight free-form crop router distills these labels into a one-shot inference policy that predicts whether to review, expected utility, and a continuous crop box from the global state. Across LLaVA-1.5-7B, InternVL2.5-8B, and Qwen2-VL-2B-Instruct, GapSight improves the Base no-zoom baseline on six benchmarks spanning OCR, documents, charts, infographics, VStarBench, and MME-RealWorld-Lite. On InternVL2.5-8B, GapSight raises the six-benchmark average from 52.25 to 64.29, above CropVLM (57.16), ViCrop (55.84), and ZoomRefine (54.43). Mechanism analyses show that the router rescues concrete wrong answers, adapts its action rate by task, and forms a favorable token-performance profile. These results position loss-gap supervision as a practical route to teaching VLMs when and where to look again.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Detail-Centric Questions
Low-Resolution Global View
Visual Tokens
Global Context
Innovation

Methods, ideas, or system contributions that make the work stand out.

visual re-reading
loss-gap supervision
free-form crop routing
selective attention
💼 Related Jobs
No related jobs found.
J
Jinchang Zhu
The Hong Kong University of Science and Technology (Guangzhou)
R
Rong Fu
University of Macau
Yi Ding
Yi Ding
University of Electronic Science and Technology of China
Deep LearningMedical Image segmentation
C
Chenghao Wu
The Hong Kong University of Science and Technology (Guangzhou)
Y
Ying Liu
The Hong Kong University of Science and Technology (Guangzhou)
Menglin Yang
Menglin Yang
HKUST(GZ) | Yale University | CUHK
Hyperbolic Representation LearningTransformerRecommender SystemLLM