Khomdet Phapatanaburi, Khwanjit Orkweha, Watcharakorn Pinthurat, Chaiwat Makhonpas, Wongsathon Pathonsuwan, Patikorn Anchuen, Monthippa Uthansakul, Peerapong Uthansakul
Knocking-sound analysis provides a non-destructive method for assessing durian ripeness; however, most convolutional neural network methods mainly emphasize magnitude information while disregarding phase information that reflects the overall signal structure. KnockNet is proposed as a model that processes the magnitude and phase feature domains and integrates them to exploit the complementary benefits of both representations. Mel-frequency cepstral coefficients contain the power distribution across frequency bands, while modified group delay cepstral coefficients capture delicate phase dynamics. Each feature stream is processed through specialized branches using a local channel attention mechanism, which emphasizes significant spectral channels, and a streamlined convolutional mixer block captures global context (broader temporal–spectral patterns), while convolutional layers with local channel attention capture local context (fine-grained spectral details). The resulting embeddings are combined to improve interpretation of the knocking event. We collected a dataset of 52 Monthong durian fruits in Chanthaburi Province, Thailand, and evaluated the model using five-fold cross-validation. KnockNet reached 96.34 ± 0.40 % accuracy (Precision/Recall/F1-score: 96.44 / 96.34 / 96.34 % ), exceeding mel-frequency cepstral coefficient-only ( 95.48 ± 0.61 % ), CNN ( 94.21 ± 0.81 % ), convolutional neural network-long short-term temory ( 94.82 ± 1.13 % ), and support vector machine ( 93.83 ± 0.83 % ). These findings confirm that the combined use of magnitude and phase information, improved by focused attention and proper mixing, produces a practical and effective instrument for assessing durian maturity in the field. • Introduces KnockNet , the first dual-stream MFCC–MGDCC network for non-destructive durian ripeness grading. • Employs Local Channel Attention and ConvMixer to fuse local and global acoustic cues. • Achieves 96.34% accuracy on a 52-fruit Monthong dataset, outperforming CNN, CNN-LSTM and classical ML baselines.