Nasser Parishad, Mehmet Yildirimoglu, Mark Hickman
Among congestion mitigation strategies, congestion pricing remains one of the most effective tools to reduce peak-hour traffic. Although previous studies have examined the fairness and equity implications of alternative pricing schemes only after the implementation of pricing schemes, it is challenging to explicitly incorporate equity considerations directly into the design of congestion pricing algorithms. This study proposes a constrained, data-driven congestion pricing framework based on reinforcement learning (RL) to maximise network outflow while maintaining equity. A trip-based three-dimensional Macroscopic Fundamental Diagram (3D-MFD) simulation model is developed to serve as the RL environment, capturing multi-modal interactions among cars, buses, and walking. Traveller heterogeneity is incorporated through variations in origin-destination pairs, trip lengths, departure times, and individual values of time (VoT). Two tolling strategies (cordon-based and distance-based) and two RL algorithms (Deep Q-Network and Deep Deterministic Policy Gradient) are implemented and compared. An equity-constrained extension based on constrained DDPG with Lagrangian relaxation is further developed. The results indicate that all RL agents significantly mitigate congestion and reduce total travel time by up to 50%. The proposed equity-constrained framework achieves comparable efficiency while distributing cost burden more equitably among travellers. Robustness and generalisation are assessed through sensitivity analyses with biased inputs in the range of [ − 20 % , 20 % ] and transferability tests under varying demand and regional capacity conditions. Across all scenarios, comparisons with a traditional feedback controller and Bayesian optimisation demonstrate that the RL-based approaches consistently outperform benchmark methods.