Hui-Chao Tian, Xiao-Hui Li
Overall evaluations were moderately positive but varied across rating items. School or period identification and formal analysis detail received relatively high ratings, whereas the bias-related item, D7, received the lowest mean and greatest dispersion. Because all items were positively keyed, lower D7 scores indicated greater perceived bias-related problems. Internal consistency, correlation, and principal component analyses showed substantial overlap among D1-D6, indicating a broad textual-quality impression, whereas D7 was less strongly integrated with this response pattern. Artwork familiarity, prior AI-based artwork-analysis experience, and art history or visual culture coursework were associated with more favorable ratings. General AI-use frequency showed no significant adjusted association.
INTRODUCTION: Generative artificial intelligence increasingly produces explanatory texts for artworks, but how students evaluate these texts remains insufficiently understood. This study examined undergraduate art students' evaluations of ChatGPT-generated art interpretations when the AI source was explicitly disclosed.
METHODS: A total of 134 students evaluated Chinese-language interpretations of 10 artworks associated with United States art history. For each artwork, students reported their familiarity and rated school or period identification, formal analysis detail, content or theme interpretation, inferential rigor, cultural or stylistic conjecture, learning inspiration, and bias-related perceived quality.
RESULTS: Overall evaluations were moderately positive but varied across rating items. School or period identification and formal analysis detail received relatively high ratings, whereas the bias-related item, D7, received the lowest mean and greatest dispersion. Because all items were positively keyed, lower D7 scores indicated greater perceived bias-related problems. Internal consistency, correlation, and principal component analyses showed substantial overlap among D1-D6, indicating a broad textual-quality impression, whereas D7 was less strongly integrated with this response pattern. Artwork familiarity, prior AI-based artwork-analysis experience, and art history or visual culture coursework were associated with more favorable ratings. General AI-use frequency showed no significant adjusted association.
DISCUSSION: Under disclosed-source conditions, the findings show how students organize perceived-quality judgments of AI-generated art interpretations. Such interpretations may support observation, discussion, and evidence-verification activities in art education.