InterSing: Explicit Interaction Dynamics for 3D Duet Singing Animation and Beyond

📅 2026-09-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出InterSing框架,通过引入交互logits和基于音频特征与交互动态的扩散模型,生成更协调且富有表现力的二重唱3D头部动画。
📝 Abstract
We present InterSing, a framework for generating realistic 3D head animations for duet singing performances. Unlike solo singing, duet performance requires each singer to balance individual expressiveness with intermittent interaction at musically salient moments, such as phrase boundaries, synchronized rhythms, and call-and-response passages. Because these interactions are sparse and rhythm-dependent, existing audio-driven animation methods and conversational interaction models do not adequately capture their structure. Our key insight is that duet coordination can be represented as a time-varying signal that reflects how strongly performers engage with one another throughout a song. Based on this observation, we introduce interaction logits, an interpretable latent representation that models the degree of cross-performer engagement at each time step. We learn these logits using weak supervision and use them to condition an interaction-aware diffusion model jointly driven by audio features and interaction dynamics. This formulation enables unified multi-mode generation, spanning independent motion, coordinated behavior, and smooth transitions between them. Experiments show that InterSing generates realistic and expressive singing head animations with stronger coordination and musical alignment than existing methods, while preserving each performer's characteristic motion style. We further demonstrate that the same formulation generalizes to multi-singer performances and provides intuitive control over when and how performers engage.
Problem

Research questions and friction points this paper is trying to address.

duet singing
interaction dynamics
audio-driven animation
realistic 3D head animations
coordinated behavior
Innovation

Methods, ideas, or system contributions that make the work stand out.

interaction logits
duet singing animation
time-varying signal
diffusion model
multi-mode generation
🔎 Similar Papers
No similar papers found.
Yihan Zhou
Yihan Zhou
Tsinghua University
ControlRobotics
Z
Zikai Huang
School of Computer Science and Engineering, South China University of Technology, Guangdong, China
Y
Yuyang Yu
School of Computer Science and Engineering, South China University of Technology, Guangdong, China
X
Xuemiao Xu
School of Computer Science and Engineering, South China University of Technology, Guangdong, China; Guangdong Engineering Center for Large Model and GenAI Technology; State Key Laboratory of Subtropical Building and Urban Science, Ministry of Education Key Laboratory of Big Data and Intelligent Robot
Cheng Xu
Cheng Xu
Beijing Key Laboratory of Information Service Engineering, Beijing Union University
Visual intelligenceInformation Security
Shengfeng He
Shengfeng He
Singapore Management University
Visual ComputingGenerative ModelsComputer VisionComputational PhotographyComputer Graphics