Pengfeng Lu, Sei-ichiro Kamata, Mengyunqiu Zhang, Weilian Zhou
Training deep neural networks (DNNs) remains challenging due to dying activations and structural loss, especially in computer vision. While complex networks with shortcut paths often outperform plain networks empirically, a clear theoretical explanation is lacking. Moreover, these networks still suffer from structural degradation. In this paper, we provide a theoretical analysis showing that neuron survival in complex DNNs is lower bounded by the probability of their shortest path length, explaining their improved trainability. We also show that downsampling causes significant structural loss and performance decline. To address these issues, we propose a Fractal Network with Wavelet Propagation (FNWP). Guided by shortest path length theory, FNWP incorporates a contraction mapping to improve neuron survival without added complexity. Its Wavelet Propagation module performs dynamic multi-scale wavelet decomposition to reduce structural loss. Furthermore, we introduce Approximation Convolution, a drop-in upgrade to standard convolutions with no extra cost. FNWP achieves 81.9% top-1 accuracy on CIFAR-100 with 14.4M parameters and 90.5% F1 score on the Kuzushiji dataset for high-resolution optical character recognition. It also delivers strong performance on general object detection tasks and time series classification benchmarks, and shows strong robustness to depth, making it a scalable architecture for deep learning.