İsmail Öner, Erdal Özbay
The rapid advancement of Large Language Models (LLMs) has significantly complicated the distinction between AI-generated and human-written texts. This challenge becomes particularly pronounced in formal and structurally constrained texts, such as academic writing. In this study, a deep learning approach based on the XLM-RoBERTa architecture is proposed for detecting deepfake (DF) texts, with a focus on achieving strong generalization capability within the academic domain. A large-scale dataset comprising 63,000 human-written and AI-generated texts (from Llama-3.1, Gemma-2, Qwen-2.5, Phi-3, Falcon, and Mistral) was constructed. The proposed multi-model data strategy is designed to encourage the model to learn structural and stylistic distinctions between human and AI-generated texts, rather than memorizing model-specific stylistic patterns, thereby reducing false positive rates, particularly for formal human-written content. To analyze the model’s learning behavior, no preprocessing was applied to the training data. The model was evaluated on two independent test sets (preprocessed and non-preprocessed), neither of which was seen during training. Experimental results show that the model achieves an F1-score of 99.76% on the validation set, while maintaining 93.42% accuracy and 94.67% recall on the preprocessed (Zero-Artifact) test set. These findings indicate that the model relies on inherent linguistic and structural patterns of AI-generated text in formal context, rather than dataset-specific superficial artifacts, suggesting improved robustness and generalization in academic integrity applications.