SADL: What to Ignore? A Benchmark for Subject-Aware Distractor Localization
Existing vision models lack subject-awareness, making it difficult to accurately identify and remove distractors in image editing without compromising scene semantic consistency. This work formalizes, for the first time, the task of Subject-Aware Distractor Localization (SADL) and introduces the first real-world benchmark for this task, comprising 1,800 cases with 14,617 annotated candidate objects. The authors propose a two-stage vision-language model (VLM) pipeline grounded in five inclusion factors and three contextual exclusion rules. Evaluation across seven VLMs reveals strong identification capabilities but exposes a systematic over-suppression bias during the exclusion phase. The SADL benchmark serves as a critical diagnostic tool for subject-conditioned reasoning in multimodal systems.