IWC-Bench: Evaluating Web Application Generation from a Software Testing Perspective

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决现有评估方法的局限性,提出IWC-Bench,通过代码覆盖率引导交互式探索生成的Web应用,并从视觉美感、可用性和需求一致性三个维度进行评价。
📝 Abstract
Human evaluation provides a direct measure of the quality of LLM-generated web applications. However, fitting human judgments through automated evaluation remains challenging. Static benchmarks can credit functionality that exists in source code but is unreachable at runtime. Interactive benchmarks exercise the application, yet incomplete exploration can cause them to miss implemented functionality and confound application defects with agent execution failures. To address these limitations, we propose IWC-Bench, an interactive benchmark for evaluating web application generation from a software testing perspective. IWC-Bench instruments each generated application and uses code coverage to guide an agent in exploring its functionality through user-simulated interactions. It then abstracts the interaction trace into a state-transition graph and evaluates the application along three dimensions: visual aesthetics, usability, and requirement alignment. By separating exploration from scoring, IWC-Bench collects runtime evidence without constraining exploration to predefined acceptance criteria. IWC-Bench comprises 369 real-world user requirements and 5,088 acceptance criteria. Evaluation of 16 frontier LLMs reveals distinct strengths across the three dimensions, with no model leading on every dimension. On 197 validated sessions sampled from an internal arena, IWC-Bench achieves 85.3\% agreement with human preferences, with agreement generally increasing as the score difference between paired applications grows. Further experiments show that coverage guidance improves exploration coverage and the model rankings remain stable when the judge model is replaced.
Problem

Research questions and friction points this paper is trying to address.

Automated Evaluation
Web Application Generation
Software Testing
Interactive Benchmarks
Code Coverage
Innovation

Methods, ideas, or system contributions that make the work stand out.

interactive benchmark
code coverage
state-transition graph
web application generation
software testing
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
C
Chenxu Liu
Hunyuan Team, Tencent
Z
Zilu Zou
Hunyuan Team, Tencent
P
Peizhong Gao
Hunyuan Team, Tencent, Tsinghua University
J
Jiawen Tao
Hunyuan Team, Tencent, Peking University
Zhexin Zhang
Zhexin Zhang
Tsinghua University, CoAI Group
NLPAI Safety & Alignment
G
Guang Chen
Hunyuan Team, Tencent
Haowei Lin
Haowei Lin
Peking University
LLMAI4Science
Y
Ying Zhou
Hunyuan Team, Tencent
Tianyi Bai
Tianyi Bai
Hong Kong University of Science and Technology(HKUST)
Large Language Models
D
Dolly Deng
Hunyuan Team, Tencent
S
Suncong Zheng
Hunyuan Team, Tencent
M
Maxm Pan
Hunyuan Team, Tencent