Zikun Meng, Wen Zhang, Jian Shi, Yan Liu, Shuo Liu, Qiankun Yu
Matched-field processing (MFP) is a conventional method for underwater acoustic target localization but is often sensitive to systematic environmental mismatch. This paper presents a transformer-based deep learning framework for joint two-dimensional (2D) range and depth localization of a fixed underwater sound source. Using the Elba-93 sea trial dataset, the model is trained on synthetic acoustic data generated by the KRAKEN propagation model and evaluated on experimental data from a 48-element vertical line array. The proposed model is compared against MFP baselines for Bartlett, minimum variance distortionless response, robust sub-array minimum variance distortionless response (SA-MVDR), and neural network baselines for one-dimensional convolutional neural network (1D-CNN) and multi-layer perceptron (MLP). Results demonstrate that the transformer achieved a 2D root mean square error of 17.59 m, which reflects the repeatability and residual bias at this fixed, on-grid location under one specific simulation-to-experiment mismatch condition. Under the same fixed-position test condition, the transformer produced lower repeated-estimate errors than the 1D-CNN and MLP baselines. Furthermore, spatial sparsity analysis reveals that the model maintains consistent accuracy using input from a single hydrophone during the testing phase. Independent evaluations indicate that the transformer encodes full-array spatial priors into its neural weights during synthetic pre-training, enabling the transfer of spatial array gain to support single sensor deployments and offering a feasible pathway for simplified underwater surveillance systems.