Weiming Zeng, Yi Sun, H.X. Shen, Zhenyu Shen, Qingqing Liu, Jianrong Lv, W. Chen, Yushan Chen, Fan Zhichun
Fiber-optic distributed acoustic sensor (DAS) offers great potential for railway event monitoring due to its high sensitivity and robustness in complex environments. However, accurate recognition of acoustic events remains challenging, under real-time constraints where only short-duration signal segments with limited discriminative information are available. To overcome this, an efficient recognition framework was proposed by integrating multi-encoding image fusion and a lightweight transformer-convolutional network (CNN). Specifically, 0.032-second DAS signal segments were converted into complementary image representations, which were then fused into RGB images to enhance feature diversity. These multimodal images were processed by a compact backbone that combined vision transformer (ViT) modules with convolutional components, enabling effective extraction of multi-scale local and global features. The hybrid framework effectively mitigated the lack of temporal information in short segments, achieving 98.19% accuracy with only 1.33 million parameters and 3.3 ms latency per sample, substantially faster than existing methods. Compared to baseline models, the proposed architecture achieved superior accuracy and compactness, making it highly suitable for real-time and edge deployment in DAS-based monitoring systems.