Haoran Zhang, Chunhui Zhao, Zheng-Guang Wu
In industrial and robotic control tasks, reference trajectories often concentrate within specific frequency bands, whereas noise and disturbances may dominate other frequency ranges. Conventional reinforcement learning (RL) methods typically employ time-domain cost functions that weight all frequency components equally, leading to excessive sensitivity to high-frequency noise and limited tracking accuracy. This limitation restricts the applicability of RL to time-varying trajectory-tracking problems. To address this issue, we propose a frequency-shaped implicit online natural actor-critic (FS-IONAC) approach that incorporates frequency-domain specifications into the RL framework. First, we theoretically show that the standard quadratic cost function inherently assigns uniform weighting across the frequency spectrum. To overcome this drawback, we introduce stable shaping filters centered on task-relevant frequencies and embed their dynamics into an augmented Markov decision process (MDP), thereby inducing a frequency-shaped objective while preserving the Markov property. Then, we develop a model-free natural actor-critic algorithm with implicit critic updates to learn the optimal tracking policy online without requiring knowledge of the system dynamics. We further prove that the proposed algorithm converges almost surely under two-timescale stochastic approximation conditions. The effectiveness of FS-IONAC is validated through illustrative examples, including a robotic control task, where it achieves improved tracking performance and sample efficiency compared with conventional baselines.