Fatima A. Hikmat, Mouayad A. Sahib
This study proposes a new intelligent framework tocope with the challenges involved with dynamic resource allocationin the 5G network environment based on Proximal PolicyOptimization (PPO), which is one of the most successful DeepReinforcement Learning (DRL) techniques. We have reformulatedresource allocation as a Markov Decision Process (MDP). Here, the"state" represents the current status of the network in terms ofdemand, interference, and channel quality. At the same time, the"Action" represents the allocation decision made for each serviceslice in terms of spectrum, capacity, and time. The proposed modelfocuses on balanced dynamic resource allocation across three mainsegments: eMBB, URLLC, and mMTC, through ensuring thatQoS requirements for each segment are met without impact to theoverall system performance. Our simulation results havedemonstrated excellent performance by the proposed algorithmwhen compared to traditional algorithms (i.e., GA, PSO, QLearning,and Round Robin). In our results, we showed athroughput increase of approximately 180 Mbps, energy efficiencyof 0.91 bps/joule, a Fairness Index of 0.88 overall performanceimprovement between 12% to 15%. As a result of the simulationresults, we believe that the PPO-MDP Framework is a good,realistic option for optimizing the use of resources within adynamically segmented environment, thus improving the ability ofa 5G system to efficiently and sustainably respond to a variety ofservice demands.