Score
Designs multimodal behavior coordination systems for avatars, producing animation controllers and synchronization mechanisms that coordinate motion, gaze, and speech across modalities.
Existing systems struggle to unify multimodal perception, embodied expression, and multi-agent collaborative decision-making within a shared physical space, limiting natural and scalable human–multi-robot interaction. This work proposes a unified framework for human–multi-agent interaction that integrates multimodal perception, large language model (LLM)-driven embodied planning, and a centralized coordination mechanism within a multi-agent architecture. The mechanism dynamically manages speaking turns and behavioral participation to effectively prevent conflicts and enable coordinated strategies across speech, gesture, gaze, and locomotion. Evaluated on a dual-humanoid robot platform, the system demonstrates robust cross-agent collaborative reasoning and embodied responsiveness, significantly enhancing the naturalness and scalability of human–robot interaction.
Current video avatar models generate physically plausible animations but are limited to low-level audio-visual synchronization, lacking deep semantic understanding of emotion, intent, and context. To address this, we propose a cognition-driven virtual character generation framework. First, we employ a multimodal large language model (MMLM) to extract high-level semantics and guide animation generation. Second, we introduce a Pseudo Last Frame mechanism to enable cross-modal coordination and conflict mitigation within a multimodal diffusion Transformer (DiT) architecture. Third, we integrate joint audio-image-text encoding to ensure semantic consistency across modalities. Experiments demonstrate state-of-the-art performance in lip-sync accuracy, video quality, motion naturalness, and textual semantic fidelity. Moreover, our framework generalizes effectively to multi-character and non-human agent scenarios.
This work addresses the challenge of explicitly controlling nonverbal behaviors—such as gaze, head motion rhythm, and emotional expression—in speech-driven facial animation during dyadic conversations. We propose the first 3D talking-head prior model that enables fine-grained, explicit control over these behaviors by disentangling their underlying signals, thereby allowing precise manipulation of a virtual avatar’s listening, reactive, and interactive actions. To facilitate training, we introduce an automated pipeline for extracting pseudo-labels of nonverbal behaviors from in-the-wild conversational videos. Our approach employs a causal flow-matching Transformer to model the influence of audio, interlocutor motion, and user-specified control signals on target head dynamics, integrated with a Gaussian head avatar prior to achieve high-fidelity, real-time animation without retraining. Experiments demonstrate superior performance over existing dyadic motion baselines in terms of motion quality, expressiveness, and diversity. Code and dataset are publicly released.
Existing methods struggle to generate controllable talking avatars that exhibit embodied, text-aligned interactions with surrounding objects, primarily due to limited environmental awareness and the trade-off between controllability and visual fidelity. To address this, this work proposes InteractAvatar, a novel dual-stream framework that achieves, for the first time, text-driven generation of embodied talking avatars capable of object interaction. The approach decouples environmental perception from video synthesis through two parallel modules: a Perception-Interaction Module (PIM) and an Audio-Interaction-aware Generation Module (AIM). Enhanced environmental understanding is achieved via object detection, while a motion-video aligner ensures consistency between generated motions and semantic content. Extensive experiments on the newly introduced GroundedInter benchmark demonstrate that InteractAvatar significantly outperforms existing methods, producing high-quality, semantically aligned interactive talking avatar videos.
This study investigates how avatar animation modalities affect remote meeting efficacy. A within-subjects experiment with 68 employees compared three avatar conditions—static image, speech-driven animation, and real-time webcam-driven facial and head pose animation—during a collaborative decision-making task. Results show that webcam-driven avatars significantly outperformed both speech-driven and static avatars in meeting effectiveness (p < 0.01), user satisfaction, and perceived inclusivity. Qualitative thematic analysis identified “holistic motion”—coordinated, contextually appropriate gestures and expressions—as a critical factor shaping user perception. Critically, this work provides the first empirical validation of the design principle that “meaningful motion outweighs photorealism,” demonstrating that semantically grounded, low-fidelity animations yield superior interpersonal outcomes. The findings establish a practical, computationally lightweight technical pathway for inclusive virtual meeting systems and offer foundational theoretical support for motion-centric avatar design in distributed collaboration.
Existing methods for digital human generation lack the capability to unify multimodal information—including text, audio, motion, and visual cues—within a single modeling framework. This work proposes a human-centric, unified multimodal autoregressive architecture that jointly models seven synchronized modalities. The approach leverages modality-specific tokenizers, a semantic video reparameterization strategy that reduces video token count by 4× while preserving dynamic details, and an “intra-modality reasoning” mechanism that decomposes cross-modal tasks into stepwise reasoning chains. Integrated with multi-task pretraining and a semantics-driven video diffusion decoder, the proposed method achieves state-of-the-art or competitive performance across diverse digital human generation tasks, significantly enhancing both output fidelity and controllability.
This work addresses the challenges of insufficient nonverbal responsiveness, coordination, and synchrony in human-robot partnered dancing by proposing a multimodal embodied language model. The approach extends the vocabulary of a large language model with discrete motion tokens and pairwise relational tokens, while incorporating audio input to enable real-time full-body salsa motion generation in response to both a human leader and background music. Key innovations include skeleton-dynamics-based automatic text descriptions for semantic alignment of tokens and a two-stage “token-to-diffusion” generation pipeline. Experimental results demonstrate that the proposed method significantly outperforms existing baselines in motion quality, musical synchronization, partner coordination, and dyadic spatial consistency, as validated through comprehensive subjective and objective evaluations.
This work addresses the challenges of reproducibility, limited adaptability, and deployment complexity commonly encountered in social robotics research by introducing M—a low-cost, open-source, and highly modular social robot platform. M integrates a modular mechanical design, multimodal perception capabilities, a streamlined expressive behavior-driven architecture, and a native ROS 2 software stack, enabling clear decoupling among perception, expression, and data management. The platform also provides a consistent simulation environment that facilitates efficient sim-to-real transfer. Through participatory design and a week-long in-home field deployment, the study demonstrates M’s practicality, scalability, and robustness in real-world settings, showcasing its effectiveness in representative interaction paradigms such as storytelling and conversational tutoring.
Existing approaches struggle to generate high-quality, coherent full-body motions from streaming audio—encompassing both speech and music—with low latency. This work proposes the first unified architecture for streaming audio-to-motion generation that operates without domain labels, continuously producing temporally coherent 3D character animations from incremental audio input. The method integrates reinforcement learning to optimize online motion quality, incorporates a large language model tool-calling interface for semantically controllable gesture synthesis, and employs an unsupervised domain generalization strategy to enhance robustness across diverse scenarios. Experimental results demonstrate that the system significantly outperforms current real-time methods in both motion fidelity and audio-motion synchronization, enabling low-latency, high-fidelity, and multi-scenario deployment of interactive virtual avatars.
This study investigates how users experience AI-generated speech and behavior in virtual reality as extensions of their own expression, raising critical questions about self-identity, agency, and authorship. To explore this, we developed ProxyMe, a VR prototype integrating embodied avatars, voice cloning, and AI-powered speech enhancement, enabling users to interact through avatars whose vocal content and delivery are modulated by AI. The work introduces the novel concept of “avatar-mediated self-extension” and systematically demonstrates how varying levels of delegation and user control influence the internalization of AI agents and the subsequent reconfiguration of perceived self-boundaries. By elucidating the mechanisms underlying self-perception in human–AI fused interactions, this research establishes a new theoretical paradigm and practical framework for designing deeply integrated AI-augmented experiences.