Hao Ni, Qian Fu, John Easton, Ning Zhao, Theodoros N. Arvanitis
Reinforcement learning (RL) has been increasingly applied to railway decision problems, particularly those that are sequential, dynamic, and difficult to optimise online. This paper presents a systematic scoping review of RL applications in railway systems published from 2006 to March 2025. A total of 162 studies were identified, screened, and analysed through a structured coding framework. The retained literature was organised into six application domains, namely Traffic Planning and Management, Automated Driving and Control, Maintenance and Inspection, Railway Communications, Passenger, and Safety. For each domain, the review examined how railway problems were formulated in RL terms, including state representation, action design, reward construction, learning environment, and validation setting. The results show that current research is concentrated in Traffic Planning and Management and Automated Driving and Control, while Passenger, Safety, and parts of Railway Communications have received less attention. RL applications are typically built through highly engineered task formulations rather than shared modelling standards, and the evidence base remains strongly dependent on simulation-based evaluation. The review further identifies persistent limitations in benchmark comparability, generalisation evidence, explainability, and deployment readiness. Overall, RL is now an established research direction in railway decision problems, but the field shows much greater diversity in problem formulation than convergence in validation practice or railway-grade deployment. Further progress will require more comparable evaluation, learning architectures that are easier to audit, stronger evidence on cross-context transfer, and closer alignment with real railway operating conditions.