SeRV: Semantic-Aligned Residual Vector Quantization for American Sign Language Generation

📅 2026-09-04
📈 Citations: 0
Influential: 0
📄 PDF
📝 Abstract
American Sign Language (ASL) generation remains challenging due to limited paired text-ASL motion data and the difficulty of learning motion representations both precise for reconstruction and predictable from linguistic input. Existing methods rely on motion tokenizers optimized for reconstruction, without explicit semantic supervision from paired text. As a result, the learned tokens remain limited in supporting semantically consistent and fine-grained ASL motion generation. To address this limitation, we propose SeRV (Semantic-Aligned Residual Vector Quantization), a semantic-aligned RVQ tokenizer for ASL generation. SeRV learns a semantically structured residual token space by combining sentence-level motion-text alignment with token-level text-conditioned supervision. Building on this tokenizer, a Hierarchical GPT predicts residual motion tokens in a coarse-to-fine manner, generating structurally coherent and semantically aligned 3D ASL motion. We further construct a large-scale reconstructed 3D ASL motion-text benchmark by recovering paired 3D motion from YouTube-ASL videos. Experiments across 375 hours of ASL video show that SeRV achieves state-of-the-art pose accuracy on both How2Sign and YouTube-ASL datasets, while producing semantically consistent 3D ASL motion directly from text.
Problem

Research questions and friction points this paper is trying to address.

American Sign Language
motion representation
semantic supervision
paired text-ASL motion data
Innovation

Methods, ideas, or system contributions that make the work stand out.

Semantic-Aligned Residual Vector Quantization
Hierarchical GPT
3D ASL Motion Generation
Motion-Text Alignment