Yuhang Yang, Yuanqing Luo, Yingyu Yang, Siqi Lv, Ziyu Wang, Jiyan Liang, Weichun Gao, Cong Geng, Yinyan Guan, Xueyong Tian, Xiangxi Kong
In waste classification, existing methods often struggle to fully integrate visual and semantic cues and lack robustness in structural modeling. In this paper, we propose an adaptive multi-modal dynamic graph neural network framework that fuses features from a 50-layer residual network (ResNet50) and the CLIP vision transformer (base/32) (CLIP ViT-B/32), adopts an adaptive dynamic graph construction mechanism (DKNN), incorporates a multi-head dynamic attention module, and leverages a supervised contrastive loss to refine the feature distribution. Comprehensive experiments on the TrashNet dataset and a proprietary dataset demonstrate that our method achieves over 99% accuracy on both benchmarks, significantly outperforming state-of-the-art approaches in classification accuracy, convergence speed, and stability. Furthermore, validation on a custom-built industrial-scale experimental platform confirms the practical applicability of the algorithm in real-world scenarios, providing an efficient and reliable solution for the industrial automation of waste sorting.