Lexara-RF: Reference-Free Metrics for Evaluating Conversational Visual Analytics Agents

📅 2026-09-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出Lexara-RF,一种无需参考基准的评估方法,通过计算一致性、意图对齐和设计有效性来评估对话式视觉分析代理的多模态输出。
📝 Abstract
Conversational visual analytics (CVA) agents powered by large language models generate visualizations and natural-language explanations from open-ended queries. Evaluating these multimodal outputs is challenging: curated reference benchmarks are costly to author, cannot comprehensively capture the space of valid responses, and are unavailable in production. Building on the Lexara evaluation framework, we introduce Lexara-RF, a reference-free set of metrics that scores CVA outputs using only the prompt, data, and model response. We reformulate evaluation as verification: 13 metrics operationalize visualization design theory and Gricean cooperative principles as computable consistency, intent-alignment, and design validity checks. On a human-rated corpus of CVA test-cases, Lexara-RF achieves alignment comparable to reference-based formulations, outperforms surface-similarity NLG baselines, and localizes structurally grounded failures with high accuracy.
Problem

Research questions and friction points this paper is trying to address.

Conversational Visual Analytics
Evaluation Metrics
Reference-Free
Large Language Models
Multimodal Outputs
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reference-Free Metrics
Conversational Visual Analytics
Gricean Cooperative Principles
🔎 Similar Papers
No similar papers found.