Yingjun Shen, Renda Shi, Kaixi Song, Yunpeng Li
DKFraudNet consistently improved downstream classification performance under weak supervision. On the independent City 2 validation set, the CatBoost instantiation achieved an accuracy of 0.911 and an F1-score of 0.909. For operational review prioritization, DKFraudNet-XGBoost required reviewing 40.2% of unverified users to capture 80% of fraud users, corresponding to a 49.71% workload reduction relative to random review, with an AUPRC of 0.954. In operator-side blind verification, genuine-member identification and false-member screening achieved accuracies of 98.9% and 93.3%, respectively.
INTRODUCTION: Fraud user identification in telecommunications is hindered by scarce, noisy, and imbalanced labels, while expert rules may provide ambiguous or contradictory evidence.
METHODS: We propose DKFraudNet, a knowledge-guided framework that integrates domain knowledge regularization, an attention-adaptive conditional generative adversarial network, and virtual category learning. Expert rules are organized into a Deterministic-Ambiguous-Contradictory evidence taxonomy for controlled pseudo-labeling. Class-conditioned augmentation alleviates data imbalance, and kernel-based similarity refines ambiguous samples. The framework was evaluated using two real-world city-level telecommunications datasets containing 59,015 and 52,183 users, respectively.
RESULTS: DKFraudNet consistently improved downstream classification performance under weak supervision. On the independent City 2 validation set, the CatBoost instantiation achieved an accuracy of 0.911 and an F1-score of 0.909. For operational review prioritization, DKFraudNet-XGBoost required reviewing 40.2% of unverified users to capture 80% of fraud users, corresponding to a 49.71% workload reduction relative to random review, with an AUPRC of 0.954. In operator-side blind verification, genuine-member identification and false-member screening achieved accuracies of 98.9% and 93.3%, respectively.
DISCUSSION: The results show that combining structured domain evidence, adaptive generative augmentation, and uncertainty-aware refinement improves robustness, data efficiency, and operational usefulness for fraud detection under weak supervision.