科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Singapore medical journal2026-09-25

Automated grading of short-answer image-based assessments using a hybrid natural language processing-AI framework: validation in radiology education.

Xu Hao Isaac Tan, Zhen Li Samantha Lee, Thida Win, Phua Hwee Tang

一句话结论 · In one sentence

The automated grading system, enhanced by NLP techniques and ChatGPT-assisted synonym expansion, demonstrated strong concordance with human grading, particularly for one- and two-word answers. This hybrid approach offers a scalable solution for short-answer assessment, with strong potential in radiology and other medical education contexts.

原始摘要(英文原文)· Original abstract
INTRODUCTION: Efficient and reliable assessment is essential in medical education. Manual grading of short-answer radiology quizzes is resource-intensive and variable. Automated marking systems can alleviate these challenges. This study introduces a novel automated grading system leveraging techniques in natural language processing (NLP) and generative artificial intelligence. Our objective was to assess the system's agreement against human grading and to analyse how response length influences grading concordance. METHODS: We obtained 840 responses from 28 radiology residents to 30 X-ray quiz questions. The system employed string-matching algorithms and rule-based logic for anatomical and laterality mismatches. A ChatGPT-created synonym dictionary was used to recognise acceptable alternative answers. Cohen's κ was used to assess inter-rater agreement across answer lengths, and logistic regression was used to examine the response length's relationship with grading disagreement. RESULTS: To establish a reference standard, two independent human raters graded the responses, achieving near-perfect inter-rater reliability (κ = 0.985). Evaluated against human consensus, the system achieved 97.5% raw agreement (816/837). Perfect agreement (κ = 1.00) was observed for one-word responses (n = 429), with near-perfect agreement for two-word responses (κ = 0.927, n = 75). Agreement decreased as response length increased, with each additional word increasing the odds of disagreement by 51.9% (odds ratio 1.519; 95% confidence interval 1.227-1.879). CONCLUSION: The automated grading system, enhanced by NLP techniques and ChatGPT-assisted synonym expansion, demonstrated strong concordance with human grading, particularly for one- and two-word answers. This hybrid approach offers a scalable solution for short-answer assessment, with strong potential in radiology and other medical education contexts.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Automated grading of short-answer image-based assessments using a hybrid natural language processing-AI framework: validation in radiology education. — 科研速览 Science Skim