Weixin Zhai, Zhuoyue He, Jinming Liu, Jiawen Pan, Caicong Wu
Agricultural machinery trajectory operation mode identification constitutes a key spatiotemporal data processing task for achieving precision agriculture. The objective of this task is to learn the features from agricultural machinery trajectories, thus assigning appropriate semantic labels to unclassified trajectory points. Existing studies have focused on spatiotemporal correlations within local areas but have failed to model dependencies at different scales, while most identification models heavily depend on the manually labelled data for training. To address these shortcomings and improve identification accuracy, we propose a dual-context encoder with spatiotemporal feature enhancement (DCST-Encoder). First, a spatiotemporal feature enhancement (STFE) scheme captures both the instantaneous motion state and the overall motion state aggregated over a temporal window of agricultural machinery, thereby enriching the expressive power of trajectory information. Next, a dual-context encoder (DC-Encoder) incorporates a global context extraction module (GCEM) and a local context aggregation module (LCAM) to comprehensively model the dependencies that exist at different scales within agricultural machinery trajectories. Finally, a dual-stream collaborative pretraining (DSCP) strategy pretrains the model, enabling the automatic learning of more generalized spatiotemporal representations from massive agricultural machinery trajectories. On paddy, wheat and corn harvester trajectory datasets, the accuracies of the DCST encoder were 91.00%, 90.32% and 90.12%, respectively, and the corresponding F1 scores were 90.99%, 90.02% and 72.53%, respectively. Compared with the current state-of-the-art models, the accuracy values were improved by 6.52%, 3.40% and 1.22%, respectively, and the F1 scores were improved by 6.58%, 7.17% and 4.10%, respectively. The proposed DCST-Encoder provides a comprehensive and generalizable framework for this task, where the limitations of local dependency modelling and reliance on labelled data are simultaneously addressed through multiscale contextual learning and self-supervised pretraining. The source code can be accessed at: https://github.com/pjw2146087/DCST-Encoder .