Yuxuan Hu, Shaohua Wang, Haojian Liang
Urban healthcare systems are fundamentally constrained by the mismatch between static resource configurations and dynamically evolving patient demand. Under the tiered healthcare system, traditional static planning methods struggle to capture the complexity and randomness of patient flows. While recent reinforcement learning (RL) approaches enable adaptive decision-making, they suffer from dimensionality explosion and unstable convergence due to massive action spaces and delayed spatiotemporal credit assignment in city-scale environments. To address this gap, we propose Twin–Game: a digital twin-driven hierarchical reinforcement learning (HRL) framework that formulates adaptive healthcare resource optimization as a “Twin Game” between a simulation-based game environment (Strategic Sandbox) and a hierarchical decision policy. First, we construct the “first twin”—an offline digital twin that serves as the Strategic Sandbox parameterized with Wuhan’s observed facility, population, and transportation data, while patient arrivals and disease profiles are generated synthetically under documented assumptions because individual-level clinical flow data are not publicly available. This environment integrates a dynamic gravity model with a two-way referral mechanism to represent the nonlinear coupling between hospital attractiveness, crowding levels, and patient choice behaviors. Second, we build the “second twin”—an Option-based HRL policy. The Manager (Macro-level Strategic Layer) uses a Deep Q-Network (DQN) for discrete spatial attention allocation; the Worker (Micro-level Execution Layer) uses Proximal Policy Optimization (PPO) for continuous, fine-grained controls such as bed expansion ratios and personnel scheduling. The two twins interact in a closed-loop game, performing strategy search and game evolution under complex constraints to optimize allocation. Experimental results from the Wuhan case indicate that the Twin–Game framework outperforms static baselines and single-layer RL in reducing average travel times, enhancing resource utilization, and improving tiered diagnosis and treatment within the simulation setting. The results should be interpreted as simulation-based decision-support evidence rather than direct clinical validation. This study provides a data-driven, game-theoretic decision support tool for building resilient urban healthcare systems.