科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Journal of Applied Psychology2026-06-18· Psychology

Scoring employment interviews with large language models: Evaluation design components, validity investigations, and best practice recommendations.

Kayden Stockdale, Louis Hickman, Siyi Liu

原始摘要(英文原文)· Original abstract
= 144). We then investigated the LLM scores' intrarater reliabilities, test-retest correlations, convergent, discriminant, and criterion evidence of validity, group differences, and measurement bias. We compared this evidence, when possible, to the same evidence for human raters and supervised machine learning models. The results suggest that ensembles of larger, newer LLMs using prompts with detailed construct information hold potential for scoring employment interviews with psychometric properties comparable to or superior to supervised machine learning models and single human raters. We detail the reasons that organizations may want to be cautious in adopting LLMs for scoring high-stakes open-ended assessments, but since organizations have already begun adopting them, we also offer best practice recommendations. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Scoring employment interviews with large language models: Evaluation design components, validity investigations, and best practice recommendations. — 科研速览 Science Skim