Institution profile

OpenNLP Lab

Academic institutionnorthamerica · us
Official website
Research library1linked papers
Opportunities0open roles
Selected work

Representative Papers

VideoVIBE: A Video-Grounded Diagnostic Benchmark for One-Shot Interactive Website Generation

Aug 10, 2026

Existing methods struggle to perform fine-grained, interpretable quality diagnosis of interactive web pages generated in a single pass from natural language, often failing to pinpoint the root causes of errors. To address this limitation, this work introduces VideoVIBE—the first diagnostic benchmark based on human interaction videos—and proposes V2Lens, a training-free multi-agent system that enables fine-grained behavioral fidelity analysis through joint visual and code verification. Integrating video question answering, multimodal large language models, multi-agent reasoning, and source-code–behavior alignment, V2Lens significantly enhances diagnostic performance across 13 state-of-the-art Video MLLMs. Using Gemini-2.5-Flash as the baseline (64.54 accuracy), V2Lens improves accuracy to 71.72, thereby substantially increasing both behavioral consistency and interpretability of generated web pages.

0 citationsRead paper
Recent publications

Latest Papers

VideoVIBE: A Video-Grounded Diagnostic Benchmark for One-Shot Interactive Website Generation

Aug 10, 2026

Existing methods struggle to perform fine-grained, interpretable quality diagnosis of interactive web pages generated in a single pass from natural language, often failing to pinpoint the root causes of errors. To address this limitation, this work introduces VideoVIBE—the first diagnostic benchmark based on human interaction videos—and proposes V2Lens, a training-free multi-agent system that enables fine-grained behavioral fidelity analysis through joint visual and code verification. Integrating video question answering, multimodal large language models, multi-agent reasoning, and source-code–behavior alignment, V2Lens significantly enhances diagnostic performance across 13 state-of-the-art Video MLLMs. Using Gemini-2.5-Flash as the baseline (64.54 accuracy), V2Lens improves accuracy to 71.72, thereby substantially increasing both behavioral consistency and interpretability of generated web pages.

0 citationsRead paper