科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Frontiers in artificial intelligence2026-01-01

A stylometric machine learning framework for authorship attribution of ChatGPT and student generated texts.

Mani P, Adaikalam Arulanandam

一句话结论 · In one sentence

The linguistic analysis revealed significant differences between the human-generated and AI-generated text categories. Among the evaluated classifiers, the SVM achieved the best classification performance, with an accuracy of 94.8%, precision of 93.7%, recall of 95.2%, and F1-score of 94.4%. Lexical diversity, sentence complexity, function-word frequency, and readability were identified as the most influential features in authorship prediction.

原始摘要(英文原文)· Original abstract
INTRODUCTION: The surge of generative artificial intelligence (AI) has created major challenges in determining authorship, evaluating student learning, and maintaining academic integrity. Distinguishing AI-generated text from human-generated writing has therefore become increasingly important for authorship verification, educational assessment, and academic honesty. METHODS: This study proposes a machine learning-based stylometric approach for authorship prediction of short-story adaptations written by non-native German language students and generated by ChatGPT. A corpus of 60 texts was created, comprising 30 student-written adaptations and 30 ChatGPT-generated adaptations based on the same source stories. The proposed framework integrates text preprocessing, stylometric feature extraction, feature selection, and machine learning-based classification. Linguistic and stylometric features included lexical diversity, word and character frequencies, sentence and paragraph length distributions, syntactic complexity, function-word usage, punctuation patterns, readability measures, and semantic consistency. Support Vector Machine (SVM), Random Forest, Decision Tree, and Naïve Bayes classifiers were evaluated. RESULTS: The linguistic analysis revealed significant differences between the human-generated and AI-generated text categories. Among the evaluated classifiers, the SVM achieved the best classification performance, with an accuracy of 94.8%, precision of 93.7%, recall of 95.2%, and F1-score of 94.4%. Lexical diversity, sentence complexity, function-word frequency, and readability were identified as the most influential features in authorship prediction. DISCUSSION: The findings demonstrate the effectiveness of the proposed stylometric framework for automated authorship attribution and AI-generated text detection. The approach has potential applications in educational assessment, plagiarism detection, forensic linguistics, authorship verification, and the development of intelligent tools for supporting academic integrity in multilingual learning environments.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

A stylometric machine learning framework for authorship attribution of ChatGPT and student generated texts. — 科研速览 Science Skim