Zarnab Kausar, Shaheryar Najam, Mohammed Alshehri, Yahya Alqahtani, Abdulmonem Alshahrani, Ahmad Jalal, Jeongmin Park
Sign language is a vital communication tool for individuals with hearing and speech impairments, yet Arabic Sign Language (ArSL) recognition remains challenging due to signer variability, occlusions, and limited benchmark datasets. To address these challenges, we propose a two-hand static and dynamic gesture recognition system that integrates keypoint-based descriptors (ORB (Oriented FAST and Rotated BRIEF), AKAZE (Accelerated-KAZE), SIFT (Scale-Invariant Feature Transform), and BRISK (Binary Robust Invariant Scalable Keypoints)) with shape-based features (smoothness, convexity, compactness, symmetry) for enhanced gesture discrimination. A distance map-based method is also used to extract fingertip keypoints by identifying local maxima from the hand centroid. An attention-enabled feature fusion strategy effectively combines these diverse features, and a long short-term memory (LSTM) network captures temporal dependencies in dynamic gestures for improved classification. Evaluated on KArSL-100, KArSL-190, and KArSL-502, the proposed system achieved 77.34%, 62.53%, and 47.58% accuracy, respectively, demonstrating its robustness in recognizing both static and dynamic ArSL gestures. These results highlight the effectiveness of combining spatial and temporal features, paving the way for more accurate and inclusive sign language recognition systems.