B. Wang, Zimeng Wang, Wenyu Zhao, Fengyuan Zhang, Wenbin Shang
Large-scale fully-meshed networks often experience prolonged routing convergence after failures because routing protocol parameters are statically configured and cannot adapt to changing network conditions. To address this limitation, this paper proposes DRL-Adapt, a deep reinforcement learning-based framework that dynamically tunes routing protocol parameters using network event-stream observations. The framework employs a dueling prioritized deep Q-network agent to jointly adjust Hello/Dead interval timers and routing weights according to real-time network states. We evaluate the proposed method using a high-fidelity simulator constructed from CAIDA topologies under diverse failure scenarios and network scales. Experimental results demonstrate that DRL-Adapt significantly accelerates recovery and reduces control overhead compared with conventional static configurations. On average, the proposed approach decreases fault convergence time by 32.7% and lowers control-plane overhead by 24.3%, while preserving routing stability. Moreover, the learned policy generalizes across networks ranging from 50 to 500 nodes and incurs minimal inference latency ($< $15 ms/decision), indicating practicality for practical deployment.