🤖 AI Summary
This study addresses the challenges of scientific figure retrieval and in-depth interpretation by proposing an "Image Scientific Concept Understanding" objective alongside a Bloom’s Taxonomy-oriented cognitive framework. Through the construction of the ALD/E-ImageMiner benchmark, comprising 1,951 expert-annotated images, this work systematically evaluates multimodal understanding, table extraction, and evidential reasoning capabilities. Furthermore, the project establishes the task framework for the ICDAR 2026 competition, thereby bridging the gap in verifiable multimodal scientific AI evaluation. Ultimately, this research provides standardized benchmarks and methodological foundations to advance intelligent scientific figure analysis from mere perception to higher-order cognition.
📝 Abstract
Scientific figures and tables encode essential experimental evidence, yet remain difficult for digital libraries and multimodal AI systems to retrieve and interpret. The ALD/E-ImageMiner benchmark and ICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching Scientific Figures provide 1,951 figures from 205 publications, expert-annotated for classification, data table extraction, summarization, and visual question answering. In these companion proceedings, we present a forward-looking perspective on how the benchmark can guide future scientific-image challenges. We examine how its tasks probe capabilities from visual and quantitative reading to domain-grounded reasoning and evidential justification, and how Bloom-informed question design can support deeper scientific understanding. We propose "scientific conceptual understanding from images" as a long-term benchmark objective, with future directions including broader domains and figure types, contextual and cross-document synthesis, hypothesis evaluation, provenance, uncertainty, counterfactual grounding, and open-ended multimodal research. This perspective connects the ICDAR 2026 challenge to a broader agenda for machine-actionable scientific visual knowledge and verifiable multimodal scientific AI.