Sign Language Recognition using Parallel Bidirectional Reservoir Computing

📅 2025-12-22
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
To address the high computational overhead of deep learning models hindering deployment on edge devices, this paper proposes a lightweight sign language recognition (SLR) system. The method leverages MediaPipe for real-time hand landmark sequence extraction and introduces a novel Parallel Bidirectional Reservoir Computing (PBRC) architecture—comprising two cooperative echo state networks—to jointly model bidirectional temporal dynamics; temporal features are subsequently fused and classified via a Softmax layer. The approach significantly reduces resource consumption and training latency: it achieves Top-1/5/10 accuracies of 60.85%, 85.86%, and 91.74% on the WLASL dataset, with training completed in only 18.67 seconds—over 170× faster than Bi-GRU. To our knowledge, this is the first SLR system enabling low-latency, low-power, and high-accuracy on-device deployment.

Technology Category

Application Category

📝 Abstract
Sign language recognition (SLR) facilitates communication between deaf and hearing communities. Deep learning based SLR models are commonly used but require extensive computational resources, making them unsuitable for deployment on edge devices. To address these limitations, we propose a lightweight SLR system that combines parallel bidirectional reservoir computing (PBRC) with MediaPipe. MediaPipe enables real-time hand tracking and precise extraction of hand joint coordinates, which serve as input features for the PBRC architecture. The proposed PBRC architecture consists of two echo state network (ESN) based bidirectional reservoir computing (BRC) modules arranged in parallel to capture temporal dependencies, thereby creating a rich feature representation for classification. We trained our PBRC-based SLR system on the Word-Level American Sign Language (WLASL) video dataset, achieving top-1, top-5, and top-10 accuracies of 60.85%, 85.86%, and 91.74%, respectively. Training time was significantly reduced to 18.67 seconds due to the intrinsic properties of reservoir computing, compared to over 55 minutes for deep learning based methods such as Bi-GRU. This approach offers a lightweight, cost-effective solution for real-time SLR on edge devices.
Problem

Research questions and friction points this paper is trying to address.

Proposes a lightweight sign language recognition system for edge devices
Utilizes parallel bidirectional reservoir computing for efficient temporal dependency capture
Achieves high accuracy with significantly reduced training time
Innovation

Methods, ideas, or system contributions that make the work stand out.

Parallel bidirectional reservoir computing for lightweight SLR
MediaPipe enables real-time hand joint coordinate extraction
Echo state network modules capture temporal dependencies efficiently
💼 Related Jobs
No related jobs found.
Nitin Kumar Singh
Nitin Kumar Singh
former NASA -JPL-Caltech
Microbial taxonomy and genomics
Arie Rachmad Syulistyo
Arie Rachmad Syulistyo
Politeknik Negeri Malang
Computer VisionMachine LearningImage Processing
Y
Yuichiro Tanaka
Graduate School of Life Science and Systems Engineering, Kyushu Institute of Technology, 2-4 Hibikino, Wakamatsu, Kitakyushu, 808-0196, Japan; Research Center for Neuromorphic AI Hardware, Kyushu Institute of Technology, 2-4 Hibikino, Wakamatsu, Kitakyushu, 808-0196, Japan
H
Hakaru Tamukoh
Graduate School of Life Science and Systems Engineering, Kyushu Institute of Technology, 2-4 Hibikino, Wakamatsu, Kitakyushu, 808-0196, Japan; Research Center for Neuromorphic AI Hardware, Kyushu Institute of Technology, 2-4 Hibikino, Wakamatsu, Kitakyushu, 808-0196, Japan