Junwei Fu, Yao Wang, Liqiang Zhu, Zujun Yu, Baoqing Guo
Abstract High-precision train self-localization is a crucial component for enabling autonomous railway operations. Conventional train localization methods heavily rely on trackside infrastructure. This paper proposes a visual virtual balise system based on visual place recognition, which replaces physical balises with camera-based visual landmarks. By storing compact global descriptors instead of raw images, the proposed system satisfies onboard storage constraints and real-time retrieval requirements, enabling infrastructure-free train localization. Furthermore, to address the strong linear structure and high visual repetition inherent in railway scenes, we propose RailVLAD, a railway-specific trainable feature aggregation network. By fusing multi-layer convolutional features, RailVLAD improves discriminative power while preserving a lightweight architecture suitable for onboard deployment. Experiments at the China National Railway Track Test Center show that the method achieves 93.68% retrieval accuracy, surpassing SIFT, NetVLAD, and VGG16 by 11.92%, 11.87%, and 6.04%, respectively. The system achieves a positioning accuracy of 1.88 m (error < 2 m), operates at 12 FPS on Jetson AGX Xavier, and can be integrated with multi-sensor localization frameworks.