IntentQA: Intent Question Answering in Videos by Cognitive Context Reasoning

📅 2026-08-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决视频中理解人类行为意图的问题,提出IntentQA任务和X-CaVIR框架,利用情境、对比和常识三种认知上下文增强模型的解释性和鲁棒性。
📝 Abstract
Video understanding requires intelligent agents to transcend mere recognition of visual facts and comprehend the underlying intents behind human actions (often termed the "dark matter" of social intelligence). To bridge the gap between visual observation and intent reasoning, we introduce a novel task, IntentQA, and contribute a large-scale VideoQA dataset specifically tailored for this purpose. However, recognizing that standard metrics may overestimate capabilities due to dataset biases, we go beyond simple accuracy to rigorously evaluate model robustness. We augment the benchmark by generating five distinct contrast sets via Large Language Models (LLMs) and introducing a "Contrast Performance Decline" metric. We propose the X-CaVIR (eXplainable Context-aware Video Intent Reasoning) framework, which leverages three types of "Cognitive Context" to enhance video analysis: i) Situational Context via a cross-modal Video Query Language (VQL) module, ii) Contrastive Context via a Contrastive Learning module, and iii) Commonsense Context via a Commonsense Reasoning module. Crucially, to overcome the opacity of traditional black-box models, we refine the integration of LLMs within X-CaVIR by employing a transparent pipeline that synergizes video captions with VQA model outputs. This approach not only improves performance by effectively utilizing rich commonsense knowledge but also renders the reasoning process explicitly interpretable. Extensive experiments demonstrate the effectiveness of our components, the superiority of X-CaVIR over state-of-the-art baselines, and its stability against perturbations on the contrast sets.
Problem

Research questions and friction points this paper is trying to address.

IntentQA
Video Understanding
Cognitive Context Reasoning
Model Robustness
Innovation

Methods, ideas, or system contributions that make the work stand out.

IntentQA
Cognitive Context
X-CaVIR
Contrast Performance Decline
Transparent Pipeline
J
Jiapeng Li
National Key Laboratory of Human-Machine Hybrid Augmented Intelligence, Xi’an Jiaotong University, Xi’an, China
Ping Wei
Ping Wei
Fudan university
Multimedia securityImage synthesis
Wenjuan Han
Wenjuan Han
Beijing Jiaotong Univerisity
Natural Language ProcessingMachine LearningArtificial IntelligenceGrammar Induction
S
Song-Chun Zhu
National Key Laboratory of General Artificial Intelligence, Beijing Institute for General Artificial Intelligence (BIGAI), Beijing, China
Lifeng Fan
Lifeng Fan
University of California, Los Angeles
Artificial IntelligenceCognitive ModelingSocial Interaction