Sign Language Recognition using Bidirectional Reservoir Computing

📅 2025-11-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Deploying sign language recognition (SLR) on resource-constrained edge devices remains challenging due to the high computational cost of existing deep learning models. Method: This paper proposes a lightweight and efficient SLR framework: it first extracts 2D hand keypoint sequences using MediaPipe; then introduces a bidirectional reservoir computing (BRC) architecture based on echo state networks (ESNs), which jointly models temporal dependencies via forward and backward reservoirs and fuses dual-path hidden states for low-complexity feature representation; finally, sequence concatenation and a lightweight classifier perform recognition. Contribution/Results: The design drastically reduces parameter count and computational load. On the WLASL dataset, it achieves 57.71% accuracy with only 9 seconds of training time—over 90% less computation than Bi-GRU—establishing a viable new paradigm for real-time, edge-deployable SLR.

Technology Category

Application Category

📝 Abstract
Sign language recognition (SLR) facilitates communication between deaf and hearing individuals. Deep learning is widely used to develop SLR-based systems; however, it is computationally intensive and requires substantial computational resources, making it unsuitable for resource-constrained devices. To address this, we propose an efficient sign language recognition system using MediaPipe and an echo state network (ESN)-based bidirectional reservoir computing (BRC) architecture. MediaPipe extracts hand joint coordinates, which serve as inputs to the ESN-based BRC architecture. The BRC processes these features in both forward and backward directions, efficiently capturing temporal dependencies. The resulting states of BRC are concatenated to form a robust representation for classification. We evaluated our method on the Word-Level American Sign Language (WLASL) video dataset, achieving a competitive accuracy of 57.71% and a significantly lower training time of only 9 seconds, in contrast to the 55 minutes and $38$ seconds required by the deep learning-based Bi-GRU approach. Consequently, the BRC-based SLR system is well-suited for edge devices.
Problem

Research questions and friction points this paper is trying to address.

Develops efficient sign language recognition for resource-constrained devices
Proposes a bidirectional reservoir computing architecture to reduce computational demands
Aims to capture temporal dependencies in sign language with minimal training time
Innovation

Methods, ideas, or system contributions that make the work stand out.

MediaPipe extracts hand joint coordinates
Bidirectional reservoir computing captures temporal dependencies
Concatenated states form robust classification representation
💼 Related Jobs
No related jobs found.
Nitin Kumar Singh
Nitin Kumar Singh
former NASA -JPL-Caltech
Microbial taxonomy and genomics
Arie Rachmad Syulistyo
Arie Rachmad Syulistyo
Politeknik Negeri Malang
Computer VisionMachine LearningImage Processing
Y
Yuichiro Tanaka
Graduate School of Life Science and Systems Engineering, Kyushu Institute of Technology, 2-4 Hibikino, Wakamatsu, Kitakyushu, 808-0196, Japan; Research Center for Neuromorphic AI Hardware, Kyushu Institute of Technology, 2-4 Hibikino, Wakamatsu, Kitakyushu, 808-0196, Japan
H
Hakaru Tamukoh
Graduate School of Life Science and Systems Engineering, Kyushu Institute of Technology, 2-4 Hibikino, Wakamatsu, Kitakyushu, 808-0196, Japan; Research Center for Neuromorphic AI Hardware, Kyushu Institute of Technology, 2-4 Hibikino, Wakamatsu, Kitakyushu, 808-0196, Japan