科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Journal of Educational Measurement2026-06-01· Rubric

Automated Coding of Communication Data Using LLM: Consistency across Subgroups

Jiangang Hao, Wenju Cui, Patrick Kyllonen, Emily Kerzabi

原始摘要(英文原文)· Original abstract
Abstract Assessing communication and collaboration at scale depends on a labor‐intensive task of coding communication data into categories according to different frameworks. Prior research has established that large language models (LLMs), particularly those from the GPT family, can be directly instructed with coding rubrics to code communication data and achieve accuracy comparable to human raters. However, whether the coding from LLMs perform consistently across different demographic groups, such as gender and race, remains unclear. To address this gap, we introduce three checks for evaluating subgroup consistency of LLM‐based coding, and empirically examine them based on data from three types of collaborative tasks. Our results show that LLM‐based coding perform consistently in the same way as human raters across gender or racial/ethnic groups, demonstrating the possibility of its use in large‐scale assessments of collaboration, communication, and other AI‐enabled conversation‐based assessments.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Automated Coding of Communication Data Using LLM: Consistency across Subgroups — 科研速览 Science Skim