Tamaththul3D: High-Fidelity 3D Saudi Sign Language Avatars from Monocular Video

šŸ“… 2026-05-06
šŸ“ˆ Citations: 0
✨ Influential: 0
šŸ“„ PDF
šŸ¤– AI Summary
This work addresses the lack of high-quality 3D annotations and dedicated reconstruction methods for Arabic sign language communities by proposing the first high-fidelity 3D avatar generation framework tailored to Saudi Sign Language. Built upon the newly introduced Ishara-500 dataset featuring high-quality SMPL-X annotations, the framework introduces Tamaththul3D—a specialized pipeline that integrates SMPLer-X for full-body pose estimation, WiLoR for refined hand reconstruction (including automatic localization and mirroring), and MediaPipe-based 2D pose supervision. Through wrist alignment via kinematic chains and a hybrid swing-twist decomposition, the method optimizes gesture articulation. Experiments demonstrate that, while preserving body pose accuracy, the approach improves hand reconstruction accuracy by up to 32% over existing methods, achieving the first complete high-fidelity 3D avatar reconstruction for Arabic sign language.
šŸ“ Abstract
Arabic Sign Language (ArSL) and its dialects serve approximately 400 million Arabic speakers worldwide, yet the community lacks high-quality 3D parametric annotations and specialized reconstruction methods for avatar generation. We address this critical gap through two key contributions: First, we introduce the first high-quality 3D parametric annotations for the Ishara-500 Saudi Sign Language dataset, providing precise SMPL-X parameters for 500 culturally authentic SSL signs. Second, we present Tamaththul3D, a specialized reconstruction pipeline designed for ArSL's unique articulation patterns. Our pipeline integrates SMPLer-X for robust body estimation, WiLoR for detailed hand refinement with automatic localization and mirroring, and MediaPipe for 2D pose supervision. Through kinematic-chain-based wrist alignment with hybrid swing-twist decomposition and 2D-supervised joint optimization, Tamaththul3D achieves state-of-the-art hand accuracy (up to 32% improvement over previous methods) while maintaining competitive body pose. Together, these 3D annotations and Tamaththul3D pipeline establish the first comprehensive framework for high-fidelity ArSL avatar reconstruction, enabling new accessibility technologies and cultural preservation efforts for the Arab Deaf community.
Problem

Research questions and friction points this paper is trying to address.

Arabic Sign Language
3D avatar
parametric annotation
monocular video
hand pose reconstruction
Innovation

Methods, ideas, or system contributions that make the work stand out.

3D sign language avatars
SMPL-X parametric modeling
hand pose refinement
monocular video reconstruction
Arabic Sign Language
E
Eyad Alghamdi
University of Jeddah, Jeddah, Saudi Arabia
S
Sattam Altuuaim
University of Jeddah, Jeddah, Saudi Arabia
O
Obay Ghulam
University of Jeddah, Jeddah, Saudi Arabia
A
Abdulrahman Qutah
University of Jeddah, Jeddah, Saudi Arabia
Y
Yousef Basoodan
University of Jeddah, Jeddah, Saudi Arabia