Adnan Nadeem, Mohammad Zubair Khan, Mehreen Sirshar, Usharani Thirunavukkarasu
Background/Objectives: Speech disorders caused by neurological, articulatory, cognitive, and behavioral impairments require accurate multimodal analysis frameworks capable of jointly understanding acoustic abnormalities and neurolinguistic inconsistencies for reliable computer-assisted speech disorder assessment. Recent multimodal deep learning approaches have utilized speech signals, linguistic transcripts, attention mechanisms, and contextual fusion strategies to improve disordered speech analysis and intelligent disorder prediction. However, existing methods frequently suffer from insufficient temporal synchronization, ineffective local-global feature learning, multimodal redundancy, poor interpretability, and reduced robustness under heterogeneous speech disorders. To address these challenges, Methods: The proposes NSX-Net (NeuroSpeech Explainable Network), an intelligent multimodal deep learning framework for speech disorder classification using Acoustic Speech Signal Data and Neurolinguistic Text/Linguistic Data. Initially, the Adaptive Speech Refinement Module (ASRM) performs noise removal, silence elimination, signal normalization, spectrogram generation, MFCC extraction, and transcript preprocessing to improve multimodal speech consistency. Subsequently, the Hierarchical Multimodal Feature Learning Unit (HMFLU) extracts discriminative feature representations through the Cross-Domain Representation Encoder (CDRE), Fine-Grained Speech Pattern Analyzer (FGSPA), and Global Sequential Dependency Learner (GSDL) for capturing both local articulation abnormalities and long-range semantic dependencies. Furthermore, the Dual-Path Attention Enhancement Block (DPAEB) emphasizes clinically important disordered speech regions using adaptive local-global attention mechanisms, while the Temporal Resolution Synchronization Module (TRSM) aligns rhythm-level and beat-level multimodal contextual structures to improve temporal consistency and suppress noise disturbances. Results: Experimental evaluation across six publicly available multimodal speech disorder datasets demonstrated that NSX-Net achieved superior performance, with 99.27% Accuracy, 99.11% Precision, 98.97% Recall, 99.02% F1-Score, and 99.41% AUC, significantly outperforming existing state-of-the-art frameworks. Conclusions: Finally, optimized multimodal contextual representations are forwarded into a Softmax classification layer for intelligent multiclass speech disorder prediction with improved interpretability and potential clinical applicability.