🤖 AI Summary
This study addresses the challenge of reliably measuring conversational states—such as cognitive load and conversational dominance—from multimodal behavioral signals, balancing predictive power, cross-task generalizability, and test–retest reliability. Leveraging the AVCAffe dataset (53 dyads across nine remote collaborative tasks), the authors construct a three-dimensional assessment framework integrating interactional, acoustic, and linguistic features, enhanced by speaker normalization to improve comparability. Findings reveal that while linguistic features exhibit the strongest predictive performance for cognitive load, they generalize poorly across tasks; acoustic features show high reliability but are strongly speaker-dependent; only interactional features—such as speaking dominance duration—robustly capture within-dyad asymmetries in cognitive load. Notably, classification of conversational dominance roles remains near chance-level across conditions. This work provides both methodological guidance and empirical grounding for selecting reliable multimodal features in social signal processing.
📝 Abstract
Measuring conversational states such as cognitive load and conversational power from multimodal behavior requires characteristic features that are not only predictive but also reliable across task contexts. We present a three-dimensional evaluation framework assessing predictive accuracy, cross-task generalizability, and test-retest reliability, applied to interactional, acoustic, and linguistic features extracted from dyadic conversations during collaborative tasks performed over a video-conferencing platform (AVCAffe dataset; 53 dyads, 9 tasks). Our results show that no single feature family dominates all three dimensions. Linguistic features show the highest predictive accuracy for cognitive load but collapse under cross-task evaluation, revealing sensitivity to task-specific vocabulary. Additionally, acoustic reliability, often reported as evidence of feature stability, degrades once speaker identity is controlled, confirming that standard prosodic features measure vocal characteristics rather than conversational state. Interaction features provide the only genuinely reliable signal, unchanged after speaker normalization. Interestingly, classifying power role remained near chance baseline across all conditions, indicating limitations of task-level aggregated behavior for predicting power role in conversation. Our findings reveal three insights: (1) linguistic features predict best but generalize poorly across task contexts; (2) acoustic reliability collapses to near-zero once speaker identity is controlled, challenging standard evaluation practice; and (3) interaction features provide the only genuinely reliable signal, with floor dominance predicting within-dyad cognitive load asymmetry. These results argue for speaker normalization and multi-dimensional evaluation as prerequisites for context-aware, robust multimodal feature selection in conversational systems.