MELLON - Multimodal Enhanced LLM for Online Navigation

📅 2026-08-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of existing web navigation agents, which are often constrained to unimodal inputs or lack effective multimodal reasoning capabilities, hindering their performance on complex tasks. Focusing on the WebShop realistic website simulation environment, the study proposes a multimodal navigation framework that integrates textual and visual information through three key innovations: the MELLON multimodal-enhanced large language model, the VQAgent architecture, and a multimodal ranker. By leveraging multimodal alignment, reasoning-based planning, and fine-tuning of large language models, the proposed approach achieves a 9.26% improvement in task-completion accuracy after only a single round of training, thereby demonstrating the efficacy and superiority of multimodal methods in web navigation.
📝 Abstract
Web navigation agents are capable of addressing various types of tasks on different websites. Current baselines on web navigation are either unimodal or lack strong reasoning abilities given multimodal inputs. Focusing on the WebShop benchmark, a real-world website simulation, we explore the alignment of text and images, as well as multimodal reasoning and planning abilities, to enhance the performance of web navigation agents. We propose three innovative multimodal enhancements: Multimodal Enhanced LLM for Online Navigation (MELLON), VQAgent, and Multimodal Ranker. MELLON demonstrates a significant improvement in task completion accuracy, with a 9.26% increase after just one epoch of training. Our findings suggest the necessity of further exploration into multimodal approaches, with a focus on more extensive training and alignment strategies to enhance the effectiveness of web navigation agents.
Problem

Research questions and friction points this paper is trying to address.

web navigation
multimodal reasoning
task completion
LLM alignment
WebShop benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

multimodal reasoning
web navigation agents
text-image alignment
LLM enhancement
task planning
🔎 Similar Papers
Ruiyu Li
Ruiyu Li
SmartMore
Computer VisionDeep Learning
H
Haoyang Cai
Carnegie Mellon University, Pittsburgh, PA, United States
Z
Zhitong Guo
Carnegie Mellon University, Pittsburgh, PA, United States
T
Tong Hu
Carnegie Mellon University, Pittsburgh, PA, United States