Yongji Guan, Shunbao Zhang, Longbang Wang, Erchao Li
ABSTRACT A central challenge in deep learning lies in achieving strong model generalisation, an area in which conventional optimisers such as stochastic gradient descent (SGD) often exhibit limitations, even though they ensure convergence. This paper introduces cascaded inertia SGD (CISGD), a novel optimisation algorithm specifically designed to address this challenge. The core mechanism of CISGD involves a hierarchical aggregation of historical gradients: it first accumulates past gradients and then compounds these accumulations to form multi‐stage inertial terms. This deep integration of gradient history allows the optimiser produce more stable optimisation trajectories, thereby encouraging convergence towards flatter minima that are empirically associated with improved generalisation. To enhance practicality and robustness, we further eliminate manual hyperparameter tuning through a lightweight feed‐forward neural network that automatically learns the optimal coefficients of the inertial terms, ensuring adaptability across diverse architectures and tasks. Extensive experiments on benchmark image classification and remote sensing change detection tasks demonstrate that CISGD consistently achieves superior generalisation performance compared with existing optimisers, demonstrating that CISGD achieves consistent, though moderate, improvements over standard optimisers while maintaining computational efficiency. This work offers a practical and theoretically grounded framework for automated and generalisable optimisation, promoting broader applications of adaptive deep learning in complex real‐world tasks.