VideoVIBE: A Video-Grounded Diagnostic Benchmark for One-Shot Interactive Website Generation

📅 2026-08-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing methods struggle to perform fine-grained, interpretable quality diagnosis of interactive web pages generated in a single pass from natural language, often failing to pinpoint the root causes of errors. To address this limitation, this work introduces VideoVIBE—the first diagnostic benchmark based on human interaction videos—and proposes V2Lens, a training-free multi-agent system that enables fine-grained behavioral fidelity analysis through joint visual and code verification. Integrating video question answering, multimodal large language models, multi-agent reasoning, and source-code–behavior alignment, V2Lens significantly enhances diagnostic performance across 13 state-of-the-art Video MLLMs. Using Gemini-2.5-Flash as the baseline (64.54 accuracy), V2Lens improves accuracy to 71.72, thereby substantially increasing both behavioral consistency and interpretability of generated web pages.
📝 Abstract
Natural-language-driven "vibe coding" enables the one-shot generation of visually rich and interactive web applications, yet reliable assessment of their quality has not kept pace. Existing evaluations often score isolated artifacts or final task outcomes, offering limited evidence about which failures occur and why. We introduce VideoVIBE, a video-grounded benchmark that transforms human-operated webpage recordings into fine-grained diagnostic tasks. It contains approximately 1.7K diagnostic Video QA instances derived from 6,338 verified failures across generated webpages, spanning semantic-logical, visual-motion, structural-temporal, and functional failures. Diagnoses are grounded primarily in recorded presentation and behavior, with webpage source code used as complementary context. We further propose V2Lens, a training-free, evidence-grounded multi-agent system that challenges and selectively refines initial video-based diagnoses through targeted visual and source-code verification. Across thirteen closed-source and open-weight Video MLLMs, Gemini-2.5-Flash is the strongest standalone model with a score of 64.54, while V2Lens reaches 71.72, an improvement of 7.18 points. Together, our results show that video-grounded evaluation can move beyond isolated artifacts and aggregate outcomes toward a behaviorally faithful and diagnostically informative account of generated application quality.
Problem

Research questions and friction points this paper is trying to address.

interactive website generation
one-shot generation
quality evaluation
diagnostic benchmark
video-grounded assessment
Innovation

Methods, ideas, or system contributions that make the work stand out.

video-grounded evaluation
one-shot website generation
diagnostic benchmark
multi-agent verification
V2Lens
🔎 Similar Papers
J
Jiajun Xu
University of Technology Sydney
Y
Yanghao Zhou
Beijing Institute of Technology
J
Jingyun Liao
Hunan University
Yu Bai
Yu Bai
Beijing Academy of Artificial Intelligence
Multi-modal ModelsEmbodied AI
J
Jinxing Zhou
OpenNLP Lab
C
Chengliang Liu
University of Macau
C
Changsen Yuan
Beijing University of Technology
B
Bo Wang
Beijing Institute of Technology
Qian Liu
Qian Liu
University of Auckland
Natural Language ProcessingSentic Computing