Yang Song, Ruoyuan Zhang
The quality of logs is crucial for the reliability of business process anomaly detection. Deployed information systems are limited by hardware and cannot generate high-quality event logs containing complex scenario information, thus requiring post-enhancement of log quality. During this process, matching the same entities across different modalities is vital. To maximize the matching degree between complex scenarios and text, we propose a two-stage method for enhancing complex scenario logs. In the first stage, cross-modal entity alignment based on image multi-feature mapping using a Vision-Language Pretraining (VLP) model is employed. Local features are extracted from images, and a gated multi-fusion module is used to dynamically fuse the global features and local multi-features of the image. The Information Noise Contrastive Estimation (InfoNCE) loss function is utilized to train and optimize the model, ensuring tight alignment between images and text. Additionally, a multi-feature enhancer is designed to match the same entities between images and text, integrating image feature information into the original logs. In the second stage, a scenario simulation graph model is constructed to effectively fuse discrete sensor data in complex scenarios with logs, supplementing the discrete data that is often overlooked in the scenarios. Finally, the performance of anomaly detection on enhanced event logs is evaluated using eight anomaly detection models on real scenarios and six commonly used event log datasets. The results show that, compared with other event log datasets, enhanced event logs have higher accuracy in business process anomaly detection. Moreover, ablation experiments are conducted on three different levels of enhanced scenario event logs to verify the effectiveness of the method.