InSight: A Benchmark for Agentic Claim Verification in Interactive Visualizations

📅 2026-09-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对动态数据分析中的证据验证问题,提出了InSight基准,通过互动可视化环境评估模型的主动求证能力。
📝 Abstract
Vision Language Models have demonstrated remarkable proficiency in interpreting static visual artifacts, but modern data analysis is inherently dynamic, requiring the active interrogation of interactive environments. Existing benchmarks are predominantly constrained to static imagery and one-shot question answering and fail to capture the epistemic demands of this domain, where evidence is frequently occluded, distributed across linked views, or conditionally revealed through user agency. In this paper, we introduce InSight, a benchmark for agentic claim verification over interactive visualizations. The dataset consists of 21,349 claims derived from human-authored analytical narratives and grounded in fully interactive web-based environments. Agents must navigate these environments to determine whether a natural language claim is supported, refuted or not verifiable given the available evidence. Unlike traditional evaluations, InSight treats interaction traces as intrinsic proxies for reasoning, enabling a rigorous audit of how models seek and synthesize visual evidence. We evaluate state-of-the-art models, revealing that interactive verification remains a non-trivial challenge. We release InSight at https://github.com/maevehutch/insight.
Problem

Research questions and friction points this paper is trying to address.

Interactive Visualizations
Agentic Claim Verification
Dynamic Data Analysis
Innovation

Methods, ideas, or system contributions that make the work stand out.

Interactive Visualization
Agentic Claim Verification
Interaction Traces