Sign Language Recognition Using Original and Synthetic Depth Image Based Point Cloud Data Models
This study addresses the limitation in existing sign language recognition research, which predominantly relies on RGB images and lacks access to real depth data, thereby hindering the application of point cloud–based methods. For the first time, it systematically compares the performance of point clouds derived from real versus synthetic depth maps for sign language recognition. Specifically, synthetic depth maps are generated from RGB images using Depth Anything V2, and point clouds are constructed from both synthetic and real depth data. These point clouds are then evaluated using representative models including PointNet variants, LSTM, and Point Gesture Map. Experimental results demonstrate that point clouds from synthetic depth consistently achieve recognition accuracy on par with or even surpassing those from real depth across multiple architectures, confirming their feasibility and practical utility as a viable alternative in scenarios where genuine depth sensors are unavailable.