Multimodal Rapport Estimation in Real-World HRI

📅 2026-08-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过在真实环境中使用多模态记录评估人机交互质量,采用零样本大语言模型及视听模型融合方法,提高了自动评估互动质量的准确性。
📝 Abstract
Evaluating interaction quality in real-world HRI is an important challenge. If interaction quality can be estimated reliably, the results can be used to improve dialogue strategies and ultimately enable robots to adapt their behavior autonomously. However, existing automatic evaluation methods have been developed primarily in controlled laboratory settings, and it remains unclear whether they can be directly applied to real-world environments, where users are free to disengage and multi-party participation may arise naturally. In this study, we investigate the automatic estimation of third-party-rated rapport scores using 62 sessions of multimodal recordings collected in a Japanese drugstore. We compare zero-shot LLMs, pretrained text, audio, and visual models, and their prediction-level fusion. The results show that, in real-world HRI, zero-shot LLMs achieve strong performance, while audio and visual models tend to provide complementary information. In particular, Gemini 2.5 Flash performs strongly as a single model, and a fusion model combining Gemini (text) with HuBERT and V-JEPA performs best overall. Further analyses showed that estimation performance varied across interaction-duration and group-size conditions. These findings suggest that rapport estimation in real-world HRI requires evaluation and model design that account for contextual variability beyond that assumed in laboratory settings.
Problem

Research questions and friction points this paper is trying to address.

real-world HRI
interaction quality
automatic evaluation
multimodal recordings
rapport scores
Innovation

Methods, ideas, or system contributions that make the work stand out.

zero-shot LLMs
multimodal fusion
real-world HRI
🔎 Similar Papers
💼 Related Jobs
No related jobs found.