AMIGO: Agentic Multi-Image Grounding Oracle Benchmark

📅 2026-03-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Current evaluations of vision-language models are largely confined to single-image, single-turn tasks, which inadequately assess their capacity for object identification and reasoning in extended, multi-image interactive settings. To address this gap, this work introduces AMIGO—the first benchmark for multi-image, multi-turn visual grounding—where an agent must locate a hidden target within a visually similar image gallery by engaging in an attribute-guided sequence of Yes/No/Unsure questions. The benchmark enforces a strict interaction protocol, incorporates a Skip penalty mechanism, and provides controllable noisy feedback, enabling comprehensive evaluation of questioning strategies, cross-turn constraint tracking, and fine-grained discriminative ability. Experiments on the Guess My Preferred Dress task systematically measure model performance across identification accuracy, evidence verification, efficiency, protocol adherence, noise robustness, and interaction trajectory quality.

Technology Category

Application Category

📝 Abstract
Agentic vision-language models increasingly act through extended interactions, but most evaluations still focus on single-image, single-turn correctness. We introduce AMIGO (Agentic Multi-Image Grounding Oracle Benchmark), a long-horizon benchmark for hidden-target identification over galleries of visually similar images. In AMIGO, the oracle privately selects a target image, and the model must recover it by asking a sequence of attribute-focused Yes/No/Unsure questions under a strict protocol that penalizes invalid actions with Skip. This setting stresses (i) question selection under uncertainty, (ii) consistent constraint tracking across turns, and (iii) fine-grained discrimination as evidence accumulates. AMIGO also supports controlled oracle imperfections to probe robustness and verification behavior under inconsistent feedback. We instantiate AMIGO with Guess My Preferred Dress task and report metrics covering both outcomes and interaction quality, including identification success, evidence verification, efficiency, protocol compliance, noise tolerance, and trajectory-level diagnostics.
Problem

Research questions and friction points this paper is trying to address.

agentic vision-language models
multi-image grounding
long-horizon interaction
hidden-target identification
visual similarity
Innovation

Methods, ideas, or system contributions that make the work stand out.

agentic vision-language models
multi-image grounding
interactive question asking
long-horizon reasoning
oracle benchmark
🔎 Similar Papers