科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Medical teacher2026-08-26

Evaluating hybrid human-LLM coding workflows for qualitative research in medical education: A generalizability study.

Emily Rush, Md Nazmul Karim, George S Yacu, Jessica N Byram, Colleen N Garnett, Nicole DeVaul, Laura Smith, Margaret Checchi, Daniel Martin, Leslie A Hoffman, Kirsten M Brown, Daniel J Mumbower, Robert M Becker, Victoria A Roach, Alison F Doubleday, Danielle N Edwards, Rebecca S Lufler, Alexandra Wactor, Sophia Boxerman, Suzanne Smith, Hannah L Herriott, Megan E Kruskie, Kyle A Robertson, Elizabeth R Agosto, Christopher Facer, Abdel Metwally, Melissa Barbosa, Dahlia Chavez, Ali Akram, Truman Steele, Seth Adler, Joshua Samaniego, Sara Aqel, Chloie Flores, Yi Gao, Emily Nguyen, Melissa Petito, Adam B Wilson

一句话结论 · In one sentence

Sequential independent LLM coding showed agreement comparable to human coding (β = -0.14, p = 0.326). Batched processing showed significantly lower human-LLM agreement (β = -0.41, p = 0.007) and substantial batch-related variance (variance = 30.16). The three LLMs produced similar agreement (joint Wald χ2 = 0.15, p = 0.930), and LLM-LLM agreement (κ = 0.750-0.755) substantially exceeded human-human agreement (κ = 0.422) and human-versus-consensus agreement (κ = 0.420-0.533). LLM disagreements were more systematically patterned across the codebook (Cramer's V = 0.554-0.585) than human disagreements (Cramer's V = 0.196). Simulations forecasted agreement ranging from 0.520 to 0.529 when two or more LLMs were paired with one to two human coders. These forecasted agreement levels for hybrid human-LLM workflows fell within the range of human-versus-consensus agreement.

原始摘要(英文原文)· Original abstract
INTRODUCTION: Large language models (LLMs) are increasingly proposed as deductive coders in qualitative research, but their measurement properties remain underexplored. This study applies generalizability theory to evaluate whether hybrid human-LLM workflow configurations can achieve reliable mode-outcomes as an alternative consensus-generating mechanism for deductive coding tasks in medical education research. METHODS: Three commercial LLMs (GPT-5.2, Claude Opus 4.5, Gemini 3-Flash Preview) coded 741 excerpts from a published audit of AI-related policy documents at 146 U.S. medical schools against a 24-subtheme deductive framework. Mixed-effects logistic regression assessed variability in agreement with human consensus across coder type (human versus LLM), excerpt characteristics (complexity and length), and coding conditions (sequential independent versus batched processing). A simulation-based D-study forecasted agreement levels for various hybrid human-LLM configurations. RESULTS: Sequential independent LLM coding showed agreement comparable to human coding (β = -0.14, p = 0.326). Batched processing showed significantly lower human-LLM agreement (β = -0.41, p = 0.007) and substantial batch-related variance (variance = 30.16). The three LLMs produced similar agreement (joint Wald χ2 = 0.15, p = 0.930), and LLM-LLM agreement (κ = 0.750-0.755) substantially exceeded human-human agreement (κ = 0.422) and human-versus-consensus agreement (κ = 0.420-0.533). LLM disagreements were more systematically patterned across the codebook (Cramer's V = 0.554-0.585) than human disagreements (Cramer's V = 0.196). Simulations forecasted agreement ranging from 0.520 to 0.529 when two or more LLMs were paired with one to two human coders. These forecasted agreement levels for hybrid human-LLM workflows fell within the range of human-versus-consensus agreement. DISCUSSION: D-study simulations support the use of hybrid human-LLM workflows to reach coding consensus for deductive reasoning tasks through mode responses, an alternative consensus-generating mechanism to traditional adjudication discussions. Future work should examine whether these patterns extend across additional deductive coding contexts and model families.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Evaluating hybrid human-LLM coding workflows for qualitative research in medical education: A generalizability study. — 科研速览 Science Skim