Does AI Understand Imaging? A Systematic Benchmark of Agentic AI for Computational Imaging Tasks
This work investigates whether multimodal agents genuinely comprehend the physical principles and inverse problems underlying computational imaging, rather than relying solely on semantic visual capabilities. To this end, we introduce ImagingBench, the first benchmark comprising 20 tasks across five major categories, and systematically evaluate state-of-the-art vision-language models—including Gemini, GPT, and Qwen—alongside specialized non-agent methods under Expert, Planner, and Forward settings. Our results reveal that current agents consistently underperform dedicated approaches across most tasks, particularly in computation-intensive domains such as lensless imaging and event camera reconstruction. Moreover, planner-guided strategies yield only marginal and unstable improvements, exposing fundamental limitations in the agents’ ability to ensure physical consistency and achieve expert-level reconstruction fidelity.