Beyond Fixed Topologies: Unregistered Training and Comprehensive Evaluation Metrics for 3D Talking Heads
Existing speech-driven 3D talking head methods are constrained by fixed mesh topologies, limiting generalization to arbitrary topologies—e.g., real-world scanned faces. To address this, we propose the first topology-agnostic speech-driven animation framework. Our method introduces: (i) a registration-free training paradigm eliminating reliance on point-to-point correspondences; (ii) a heat-diffusion-based feature prediction mechanism ensuring topology-robust geometric modeling across diverse meshes; (iii) an adaptive graph neural network that learns dynamic graph structures per input; and (iv) a multi-granularity lip-sync evaluation metric suite addressing shortcomings of conventional metrics in temporal alignment and semantic consistency. Experiments demonstrate high-fidelity animation on arbitrary-topology 3D faces—including unseen scanned data—outperforming fixed-topology baselines. We establish the first topology-independent benchmark for 3D talking heads.