Yating Yang, Bin Wang, Jun Li, Sun Pengyu, Jianfeng Wang, Weihua Li
Efficient and intuitive control of high-DoF robotic systems remains a core challenge in human-robot interaction, limiting the practical application of mobile dual-arm teleoperation. To address this issue, a vision-based human-pose-driven teleoperation system is proposed. The enhanced pose detection model is designed to improve keypoint tracking stability under occlusion. A multi-source visual fusion pipeline is employed to obtain stable human-motion trajectories and reduce pose estimation errors. A Hierarchical Pose-to-Control Mapping (HPCM) framework is further introduced to map human pose to both the dual-arm manipulators and the mobile base, aiming to achieve intuitive teleoperation of a mobile dual-arm robot. Experiments on the COCO dataset show that the Teleoperation-oriented Pose Fusion YOLO model (TPF-YOLO) achieves 2.48% higher accuracy and 0.71% better generalization than the baseline YOLOv8n-Pose model. In real-robot teleoperation experiments, the baseline YOLOv8n-Pose interface required 81.3 s on average, whereas the proposed system reduced the completion time to 62.3 s and the final placement error from within ±4.5 cm to within ±3 cm. These results demonstrate improved task efficiency and object placement accuracy in mobile dual-arm teleoperation.