科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Applied Sciences2026-05-19· Computer science

Cross-Model Deepfake Text Detection with XLM-RoBERTa: A Strongly Generalizable Multi-LLM Training Strategy

İsmail Öner, Erdal Özbay

原始摘要(英文原文)· Original abstract
The rapid advancement of Large Language Models (LLMs) has significantly complicated the distinction between AI-generated and human-written texts. This challenge becomes particularly pronounced in formal and structurally constrained texts, such as academic writing. In this study, a deep learning approach based on the XLM-RoBERTa architecture is proposed for detecting deepfake (DF) texts, with a focus on achieving strong generalization capability within the academic domain. A large-scale dataset comprising 63,000 human-written and AI-generated texts (from Llama-3.1, Gemma-2, Qwen-2.5, Phi-3, Falcon, and Mistral) was constructed. The proposed multi-model data strategy is designed to encourage the model to learn structural and stylistic distinctions between human and AI-generated texts, rather than memorizing model-specific stylistic patterns, thereby reducing false positive rates, particularly for formal human-written content. To analyze the model’s learning behavior, no preprocessing was applied to the training data. The model was evaluated on two independent test sets (preprocessed and non-preprocessed), neither of which was seen during training. Experimental results show that the model achieves an F1-score of 99.76% on the validation set, while maintaining 93.42% accuracy and 94.67% recall on the preprocessed (Zero-Artifact) test set. These findings indicate that the model relies on inherent linguistic and structural patterns of AI-generated text in formal context, rather than dataset-specific superficial artifacts, suggesting improved robustness and generalization in academic integrity applications.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Cross-Model Deepfake Text Detection with XLM-RoBERTa: A Strongly Generalizable Multi-LLM Training Strategy — 科研速览 Science Skim