Social Intuition vs. Machine Reasoning: Anticipating Human-Robot Interaction from multiple modalities

📅 2026-09-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过比较人类和不同模型在预测人机交互意图上的表现,发现即使是最先进的视觉-语言模型也难以匹敌人类的社会直觉。
📝 Abstract
Anticipating whether a person will interact from one's own perspective is a highly intuitive task for humans, that relies on a combination of cues. We investigate how humans perform at predicting a person's intention to interact from a service robot's point of view, using pose-only or full video input, then benchmark different lightweight pose-based models and state-of-the-art vision-language models. We conducted our benchmark on the HUI360 dataset on a fixed pilot subset of 100 test tracks (25 positive, 75 negative). We found that with pose-only input, human annotators outperform lightweight trained pose models but not by large margins (+0.08 in F1-Score). But when given full egocentric video with a target bounding box, human annotators perform substantially better and largely outperform the Vision-Language Models (+0.2 in F1-Score). We also compared VLMs of different size and under different input conditions, and found that the best results do not correlate with model size. Our result confirms that predicting interactions is a challenging task for social robots and that reasoning-capable models are necessary but their actual reasoning capabilities alone do not suffice to match the social intuition of humans.
Problem

Research questions and friction points this paper is trying to address.

Human-Robot Interaction
Pose-based Models
Vision-Language Models
Social Intuition
Innovation

Methods, ideas, or system contributions that make the work stand out.

pose-based models
vision-language models
human-robot interaction
social intuition
multi-modal reasoning
💼 Related Jobs
No related jobs found.
R
Raphael Lorenzo-Louis
Inria, CNRS, UL, Loria, HUCEBOT; F-54600 Villers-les-Nancy France; CEA, List, F-91120, Palaiseau, France
B
Bertrand Luvison
Université Paris-Saclay; CEA, List, F-91120, Palaiseau, France
Serena Ivaldi
Serena Ivaldi
Senior Research Scientist (Directrice de Recherche), INRIA
roboticshumanoid roboticshuman-robot interactionsocial roboticsrobot learning