When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning
This work addresses the semantic channel gap in multimodal large language models, which struggle to effectively leverage visual text for reasoning about task instructions embedded in images—such as those found in screenshots. The study is the first to explicitly identify and quantify this gap, introducing an end-to-end prompt-region grounding method that operates without OCR or region-level metadata. By embedding task instructions directly into images, the authors construct a Visualized Task Semantics (VTS) benchmark and align question-relevant regions with their semantic meanings through masked image modeling and typed semantic representations, thereby recovering clean visual features from occluded views. Evaluated across four benchmarks, the proposed approach improves VTS accuracy from 58.0 to 66.3 (+8.3 percentage points) while preserving performance on original text-interface tasks.