Xinglin Lian, Yu Zheng, Yan Liu, Fan Zhou, Chunlei Peng, Xinbo Gao
Network traffic anomaly detection is critical for cybersecurity but faces challenges in accurately identifying malicious activities. Recent zero-positive approaches, which use only normal training data under the reconstruction paradigm, have shown progress. However, encrypted network traffic obscures normal–anomalous distinctions, causing confused modeling. In addition, the “identical shortcut” problem, where models reconstruct any input with similar fidelity, produces suboptimal representations and indistinguishable detection. To address these limitations, this paper introduces ConMD, a novel Contextual Masking Knowledge Distillation framework. ConMD features distillation paradigm for discriminative representations and then pursues two objectives: effective contextual information modeling and a comprehensive anomaly metric. Specifically, we introduce context-aware local-global attention mechanisms for the student network's backbone, which capture both intra-packet and inter-packet dependencies. Additionally, a context-enhanced masking training strategy is designed to facilitate contextual interactions in normal flows. Given the structural characteristics of network traffic, we also present a new anomaly scoring with multi-view awareness, which perceive comprehensive traffic patterns. ConMD combines insights from both packet- and flow-level views to highlight deviations in anomalous network flows, thereby improving detection accuracy. Extensive experiments on three real-world datasets validate the effectiveness of ConMD, yielding consistent improvements over state-of-the-art baselines, achieving up to 2.8% and 5.1% AUC gains on the DataCon2020 and CIC-IDS2017 datasets, respectively. Our model code will be released at https://github.com/ikun0124/ConMD.