AMIGO: Agentic Multi-Image Grounding Oracle Benchmark
Current evaluations of vision-language models are largely confined to single-image, single-turn tasks, which inadequately assess their capacity for object identification and reasoning in extended, multi-image interactive settings. To address this gap, this work introduces AMIGO—the first benchmark for multi-image, multi-turn visual grounding—where an agent must locate a hidden target within a visually similar image gallery by engaging in an attribute-guided sequence of Yes/No/Unsure questions. The benchmark enforces a strict interaction protocol, incorporates a Skip penalty mechanism, and provides controllable noisy feedback, enabling comprehensive evaluation of questioning strategies, cross-turn constraint tracking, and fine-grained discriminative ability. Experiments on the Guess My Preferred Dress task systematically measure model performance across identification accuracy, evidence verification, efficiency, protocol adherence, noise robustness, and interaction trajectory quality.