Xin Wang, Yidan Su, Yimeng Fan, Wei Zhang, Mingyang Li
Cross-view geo-localization (CVGL) between unmanned aerial vehicle (UAV) imagery and satellite imagery is a key technique for autonomous UAV navigation in Global Navigation Satellite System (GNSS)-denied environments. However, most existing methods rely on energy-intensive Artificial Neural Networks (ANNs), making them difficult to deploy on resource-constrained edge computing platforms. Spiking Neural Networks (SNNs) provide a promising alternative for energy-efficient inference, but their application to CVGL still faces two challenges that remain insufficiently addressed. First, the isotropic computation used by existing SNN backbones is mismatched with the directional characteristics of spike activations. Spike activations tend to form oriented aggregation patterns along elongated geographic structures, and isotropic computation can therefore dilute directional signals. Second, the limited representational capacity of SNNs further increases the sensitivity during training optimization. However, the standard triplet loss adopts a static weighting strategy and assigns the same weight to all triplets that violate the margin constraint, which is unfavorable for learning from hard negatives. To address these challenges, we propose a framework with two core contributions. At the feature extraction level, the Directional Adaptive Convolution Module (DACM) processes spike feature maps by sequentially performing horizontal strip convolution and vertical strip convolution, thereby capturing a more complete geometric structure of directional spike clusters. At the training supervision level, we propose a Dual-dimensional Progressive Reweighting (DPR) loss, which jointly characterizes sample difficulty from pairwise difficulty and positive-pair quality difficulty. A learnable fusion parameter is used to adaptively balance these two types of difficulty information. Experimental results on the University-1652 and SUES-200 benchmarks show that the proposed framework, when equipped with the same representation learning head as its ANN counterparts, achieves competitive and, in many settings, superior performance. In terms of energy efficiency, its estimated theoretical energy consumption is over 8.8× lower than that of published ANN methods under their original configurations. Under a more rigorous matched ANN control that shares the identical architecture, the estimated energy is reduced from 29.84 mJ to 6.36 mJ, an approximately 4.7× reduction obtained at a cost of only 2.29 percentage points in R@1.