EvoGenUI-Bench: Evaluating LLMs as Multi-Turn Generative UI Assistants

📅 2026-08-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过EvoGenUI-Bench评估大型语言模型在多轮生成用户界面时的表现,使用多种方法包括执行生成的工件、截图等,发现即使最强模型也面临同步性和状态保持的问题。
📝 Abstract
Large language models can generate interactive web interfaces, but reliable generative UI requires maintaining an executable artifact as user requests evolve. We introduce EvoGenUI-Bench, a benchmark for multi-turn interface maintenance comprising 150 five-turn tasks and 750 turns across three scenarios: information presentation, executable interaction, and tool-grounded external state. We execute generated artifacts in a browser and evaluate them using screenshots, source and DOM evidence, actor traces, and runtime logs. Beyond turn-level and episode-level success, we measure cross-turn retention with Adjacent Pass Retention. Across eight models, even the strongest achieves 74.9% Turn Pass while completing only 37.3% of five-turn episodes; APR further falls to 52.4% on tool-grounded tasks. Diagnostic analysis shows that presentation failures center on information architecture, interaction failures on derived-state propagation and affordance binding, and tool-grounded failures additionally involve external-state grounding and requirement decomposition. These results reframe generative UI evaluation from judging isolated outputs to testing whether interface behavior, derived state, external state, and assistant claims remain synchronized as the artifact evolves.
Problem

Research questions and friction points this paper is trying to address.

LLMs
Generative UI
Multi-turn Interaction
Interface Maintenance
Innovation

Methods, ideas, or system contributions that make the work stand out.

EvoGenUI-Bench
multi-turn interface maintenance
executable artifact
Adjacent Pass Retention (APR)
synchronized evolution
Yue Peng
Yue Peng
University of Science and Technology of China
geometry optimizationphysical simulation
L
Lanke Xia
New York University Shanghai
Zihan Wang
Zihan Wang
New York University
Machine Learning
J
Jiahao Ye
New York University Shanghai
K
Ke Ning
New York University Shanghai
H
Hongyi Wen
New York University Shanghai