The GENEA Challenge 2026: A Large-Scale Disentangled Evaluation of Speech-Driven Gesture Generation on the Seamless Interaction Dataset

📅 2026-08-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study systematically evaluates the performance bottlenecks of speech-driven gesture generation systems across four key dimensions: motion quality, speech alignment, interactive responsiveness, and semantic expressiveness. Leveraging the Seamless Interaction dataset, the authors introduce a decoupled evaluation framework grounded in four large-scale user studies involving 869 participants and over 23,000 ratings. For the first time, they incorporate a dialogue mismatch experiment and a semantic gesture–text matching test—based on the Grounded Gestures subset—to enable independent assessment of interactivity and semantic fidelity. The results demonstrate that current systems significantly underperform human motion capture data across all evaluated dimensions, revealing critical limitations in existing approaches to generating naturalistic, contextually appropriate co-speech gestures.
📝 Abstract
This preprint presents the results of the fourth GENEA Challenge, a large-scale human evaluation of five speech-driven gesture-generation systems trained by participating teams on the Seamless Interaction dataset of dyadic conversations. As in the 2023 GENEA Challenge, we used a disentangled evaluation methodology to assess motion quality and speech alignment without confounding between the two, and performed a dyadic mismatching study to isolate the effect of listening and reacting to the interlocutor. We additionally introduce a new semantic gesture-generation task and a text-mismatching evaluation methodology using the Grounded Gestures subset of the data. In total, we ran four large-scale user studies, collecting over 23,000 votes from 869 test-takers. In the motion-realism study, the dataset's filtered segments had substantially higher motion quality than all challenge submissions (68-95% pairwise winrate). In the speech-alignment study, the motion-capture segments provided a conceptual ceiling at 62% alignment score, with the top submission significantly behind at 32% and the rest only slightly above the 0% expected of an input-independent system. In the dyadic study, motion capture again set the ceiling at 65% appropriateness score, but no submission scored substantially above chance, indicating that the systems could not yet respond to the interlocutor. Finally, the semantic mismatching evaluation found highly expressive gestures in the dataset (test-takers identified the matching transcript 79% of the time), yet almost all submissions failed to generate semantically expressive motion, with the best achieving only an 8% appropriateness score. The collected votes and outputs will be made publicly available at https://genea-workshop.github.io/2026/challenge/ to facilitate reproducibility and further research.
Problem

Research questions and friction points this paper is trying to address.

speech-driven gesture generation
disentangled evaluation
motion realism
semantic expressiveness
dyadic interaction
Innovation

Methods, ideas, or system contributions that make the work stand out.

disentangled evaluation
semantic gesture generation
dyadic mismatching
speech-driven gesture generation
text-mismatching evaluation
🔎 Similar Papers
No similar papers found.