EgoArgus: Benchmarking VLMs as Situational Assistants for Modality-Grounded User Supports

📅 2026-08-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过构建EgoArgus数据集评估视觉语言模型在日常对话视频场景中的理解和决策能力,旨在解决模型如何权衡视觉与用户语言信息的问题。
📝 Abstract
VLMs are increasingly positioned as daily assistants that perceive first-person environments, follow user dialogue, and decide how to help. Existing egocentric benchmarks mainly evaluate visual understanding in isolation, leaving open whether models can arbitrate between visual evidence and user-provided language when the two are helpful, irrelevant, or conflicting. We introduce EgoArgus, a human-annotated dataset for evaluating egocentric assistants on understanding and decision tasks in five dialogue-video daily scenarios. Our results demonstrate that it is still challenging for current VLMs as reliable egocentric assistants, which requires identifying which modality is trustworthy and deciding when intervention is warranted. Deeper analysis also shows that existing modality bias mitigation methods are quite restricted to enhance performance, providing insights to aid practioners into the deployment of current VLMs as daily assistants.
Problem

Research questions and friction points this paper is trying to address.

VLMs
Egocentric Assistants
Modality-Grounded User Supports
Visual and Language Integration
Decision Making
Innovation

Methods, ideas, or system contributions that make the work stand out.

EgoArgus
Visual-Language Models
Modality Arbitration
Bias Mitigation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.