Y Zhong, Zhiming Gui, Zhenji Gao, Xinyu Wang, Jiawen Wei
Vehicle trajectory prediction is a pivotal technology in intelligent transportation systems. Existing methods encounter challenges in effectively modeling lane topology and dynamic interaction relationships in complex traffic scenarios, limiting prediction accuracy and reliability. This paper presents Lane Interaction Transformer (LITransformer), a lane-informed trajectory prediction framework that builds on spatio–temporal graph attention networks and Transformer-based global aggregation. Rather than introducing entirely new network primitives, LITransformer focuses on two design aspects: (i) a lane topology encoder that fuses geometric and semantic lane features via direction-sensitive, multi-scale dilated graph convolutions, converting vectorized lane data into rich topology-aware representations; and (ii) an Interaction-Aware Graph Attention mechanism (IAGAT) that explicitly models four types of interactions between vehicles and lane infrastructure (V2V, V2N, N2V, N2N), with gating-based fusion of structured road constraints and dynamic spatio–temporal features. The overall architecture employs a Transformer module to aggregate global scene context and a multi-modal decoding head to generate diverse trajectory hypotheses with confidence estimation. Extensive experiments on the Argoverse dataset show that LITransformer achieves a minADE of 0.76 and a minFDE of 1.20, and significantly outperforms representative baselines such as LaneGCN and HiVT. These results demonstrate that explicitly incorporating lane topology and interaction-aware spatio-temporal modeling can significantly improve the accuracy and reliability of vehicle trajectory prediction in complex real-world traffic scenarios.