How Well Do Models Follow Visual Instructions? VIBE: A Systematic Benchmark for Visual Instruction-Driven Image Editing
This work addresses the limitation of existing image editing models and evaluation benchmarks, which predominantly rely on textual instructions and struggle to support visual directives—such as sketches—that are integral to human multimodal interaction. To bridge this gap, we introduce VIBE, the first systematic benchmark for vision-instructed image editing, which defines a three-tiered hierarchy of task complexity ranging from referential localization and shape manipulation to causal reasoning. We further develop a fine-grained automatic evaluation framework based on multimodal large language models (LMMs) and use it to assess 17 open- and closed-source models across diverse visual instructions. Our evaluation reveals that closed-source models exhibit初步 stronger instruction-following capabilities, yet all models suffer significant performance degradation on higher-order tasks, highlighting critical limitations and pointing toward promising directions for future research.