Li Gao, Qingchun Feng, Shiqi Chen, Zhijie Yang, Fengcui Fan, Lin Chen, Chunjiang Zhao
With the intensification of global agricultural labor shortage and scaled development of facility agriculture, autonomous precision harvesting robots for unstructured greenhouse environments have become an urgent need. For cluster-picking crops such as tomatoes, visual servoing enables real-time closed-loop control of the end-effector pose, addressing challenges of random fruit distribution and variable stem orientations. However, existing methods struggle to balance constraint handling with real-time efficiency. This paper proposes an MPC-Guided Reinforcement Learning visual servoing framework, innovatively combining the planning capability of optimal control with the adaptive learning ability and real-time inference advantages of reinforcement learning. The approach adopts a teacher–student paradigm: expert trajectories from the MPC controller warm-start the reinforcement learning policy through behavior cloning, followed by PPO-based fine-tuning with adaptive gain regulation and stagnation-enhanced exploration mechanisms. Simulation experiments demonstrate a 95% success rate with average positioning and orientation errors of 13.6 mm and 0.009 rad respectively. Compared to MPC baseline, task steps are reduced by 53.4%; compared to Standard PPO, success rate improves by 6%. Greenhouse field validation achieves 85.3% picking success rate and 5.63 s per fruit operation time, confirming the framework’s excellent balance among control precision, robustness, and efficiency for high-precision robotic harvesting in unstructured agricultural environments.