How Reliable Are Multimodal Signals of Conversational State? Evidence from Remote Dyadic Collaborative Tasks
This study addresses the challenge of reliably measuring conversational states—such as cognitive load and conversational dominance—from multimodal behavioral signals, balancing predictive power, cross-task generalizability, and test–retest reliability. Leveraging the AVCAffe dataset (53 dyads across nine remote collaborative tasks), the authors construct a three-dimensional assessment framework integrating interactional, acoustic, and linguistic features, enhanced by speaker normalization to improve comparability. Findings reveal that while linguistic features exhibit the strongest predictive performance for cognitive load, they generalize poorly across tasks; acoustic features show high reliability but are strongly speaker-dependent; only interactional features—such as speaking dominance duration—robustly capture within-dyad asymmetries in cognitive load. Notably, classification of conversational dominance roles remains near chance-level across conditions. This work provides both methodological guidance and empirical grounding for selecting reliable multimodal features in social signal processing.