Shengrong Lin, Yonggen Zhu, Can Cui, Lin Chen, Minfeng Tong, Jianming Wen, Kang Chen
MRI-guided ultrasound ablation has emerged as a promising precision therapy for prostate disease, using real-time temperature feedback to control thermal delivery. In this simulation-based study, we develop a controller for such systems, emulating the temperature feedback with a bioheat model. Existing control algorithms struggle to maintain ablation accuracy under dynamic tumor-boundary shifts and heterogeneous thermal diffusion. To address this, we introduce a controller that combines imitation learning with deep reinforcement learning (IL-DRL) to improve targeting precision and adaptive responsiveness during treatment. In our simulations, the IL-DRL controller achieves accelerated convergence and superior performance. When integrated with a multi-frequency ultrasound system that uses temperature feedback, the IL-DRL controller autonomously optimizes ultrasound power and rotational speed in real time, while frequency is selected by a predefined rule based on the distance to the target boundary. The proposed IL-DRL controller delivers better conformal ablation than conventional methods. In single-slice evaluations, over 97% of the ablation boundary lies within ±1 mm of the target, compared to less than 80% for conventional binary and PID controllers. For whole-gland ablation, BC + PPO1 achieves 88% precision (versus 82% for the conventional approach), while BC + PPO2 reaches a statistically comparable 85.5% (p = 0.31) but reduces treatment time by approximately 30%. The advantage of BC + PPO2 therefore lies in improved treatment efficiency rather than improved accuracy. The controller remains robust across prostates with varying acoustic properties and performs consistently in tissue-phantom validation. These results suggest the potential of the proposed controller for future clinical translation, pending dedicated validation under realistic MRI conditions and in vivo studies.