Shuwen Liang, Hailong Liu, Ping Liu, Dong Zhang, Rong Yu, Xiaowei Zhou, Zhize Zhou
This paper presents a novel approach for estimating 3D human pose from monocular video sequences in underwater scenarios, tackling the unique challenges posed by water refraction, body occlusion, low image quality and illumination distortion in underwater environments. Leveraging both 2D keypoint extraction and parametric model estimation, our method operates in a two-stage framework including preprocessing and optimization. In the preprocessing stage, a Part Attention Regressor (PARE) is adopted to dynamically estimate SMPL human body parameters, particularly adept at handling occlusions common in underwater scenarios. Additionally, a 2D keypoint detector, employing YOLO for bounding box detection and HRNet for keypoint regression, enhances feature extraction despite underwater image challenges. In the optimization stage, we propose an underwater variational autoencoder (UW-VAE), which adopts a data-driven strategy to learn the biomechanical prior distribution of underwater human poses and implicitly correct unreasonable pose parameters caused by refraction and occlusion. The optimization process incorporates constraints aligning final SMPL models with detected 2D keypoints, minimizing disparity between adjusted and original SMPL models, and ensuring temporal consistency. Furthermore, to address the scarcity of annotated underwater datasets, we build a full pipeline to generate synthetic underwater datasets with complete annotations based on UW-VAE. Experimental results on the SwimXYZ synthetic dataset show that our method achieves 51.60% PCK@0.2 and 80.13% PCK@0.5, outperforming state-of-the-art land-based methods across most stroke categories. Validation on real-world underwater swimming datasets demonstrates improved 2D keypoint accuracy after synthetic-data fine-tuning, which provides a new solution for 3D human motion analysis in underwater sports, biomechanical research and swimming training.