Mohamed Amine Hechmi, Sonia Ben Rejeb, Nidal Nasser, Sami Tabbane
The exponential growth of network traffic and connected devices poses significant challenges for load balancing in Software-Defined Networking (SDN) architectures, especially in dynamic and dense 6G vehicular networks. Traditional load-balancing techniques struggle to maintain Quality of Service (QoS) due to rapidly fluctuating traffic patterns and mobility induced variations. This paper proposes a novel andoriginaldistributed load-balancing framework combining Multi-Agent Reinforcement Learning (MARL) with Proximal Policy Optimization (PPO), which differs from existing MARL-PPO approaches by introducing a reward design that simultaneously accounts for latency, load variance, and energy efficiency. Each agent dynamically adjusts its policies based on real-time network states generated from a synthetic random traffic flow in Colab, providing a controlled and reproducible evaluation environment to optimize resource allocation and reduce controller overload. We evaluate the MARL-PPO framework against a benchmark Deep Reinforcement Learning method using Deep Deterministic Policy Gradient (DDPG). Experiments in realistic Vehicle-to-Vehicle (V2V) communication scenarios demonstrate that MARL-PPO achieves lower latency, better load distribution, and improved scalability compared to DRL-DDPG, while maintaining stable learning under variations of reward parameters. These results highlight the potential of the proposed approach to effectively address the challenges of load management in 6G enabled vehicular networks, providing reliable, low-latency, and energy-efficient communication even under highly dynamic and dense vehicular conditions.