Emily Rush, Md Nazmul Karim, George S Yacu, Jessica N Byram, Colleen N Garnett, Nicole DeVaul, Laura Smith, Margaret Checchi, Daniel Martin, Leslie A Hoffman, Kirsten M Brown, Daniel J Mumbower, Robert M Becker, Victoria A Roach, Alison F Doubleday, Danielle N Edwards, Rebecca S Lufler, Alexandra Wactor, Sophia Boxerman, Suzanne Smith, Hannah L Herriott, Megan E Kruskie, Kyle A Robertson, Elizabeth R Agosto, Christopher Facer, Abdel Metwally, Melissa Barbosa, Dahlia Chavez, Ali Akram, Truman Steele, Seth Adler, Joshua Samaniego, Sara Aqel, Chloie Flores, Yi Gao, Emily Nguyen, Melissa Petito, Adam B Wilson
Sequential independent LLM coding showed agreement comparable to human coding (β = -0.14, p = 0.326). Batched processing showed significantly lower human-LLM agreement (β = -0.41, p = 0.007) and substantial batch-related variance (variance = 30.16). The three LLMs produced similar agreement (joint Wald χ2 = 0.15, p = 0.930), and LLM-LLM agreement (κ = 0.750-0.755) substantially exceeded human-human agreement (κ = 0.422) and human-versus-consensus agreement (κ = 0.420-0.533). LLM disagreements were more systematically patterned across the codebook (Cramer's V = 0.554-0.585) than human disagreements (Cramer's V = 0.196). Simulations forecasted agreement ranging from 0.520 to 0.529 when two or more LLMs were paired with one to two human coders. These forecasted agreement levels for hybrid human-LLM workflows fell within the range of human-versus-consensus agreement.
INTRODUCTION: Large language models (LLMs) are increasingly proposed as deductive coders in qualitative research, but their measurement properties remain underexplored. This study applies generalizability theory to evaluate whether hybrid human-LLM workflow configurations can achieve reliable mode-outcomes as an alternative consensus-generating mechanism for deductive coding tasks in medical education research.
METHODS: Three commercial LLMs (GPT-5.2, Claude Opus 4.5, Gemini 3-Flash Preview) coded 741 excerpts from a published audit of AI-related policy documents at 146 U.S. medical schools against a 24-subtheme deductive framework. Mixed-effects logistic regression assessed variability in agreement with human consensus across coder type (human versus LLM), excerpt characteristics (complexity and length), and coding conditions (sequential independent versus batched processing). A simulation-based D-study forecasted agreement levels for various hybrid human-LLM configurations.
RESULTS: Sequential independent LLM coding showed agreement comparable to human coding (β = -0.14, p = 0.326). Batched processing showed significantly lower human-LLM agreement (β = -0.41, p = 0.007) and substantial batch-related variance (variance = 30.16). The three LLMs produced similar agreement (joint Wald χ2 = 0.15, p = 0.930), and LLM-LLM agreement (κ = 0.750-0.755) substantially exceeded human-human agreement (κ = 0.422) and human-versus-consensus agreement (κ = 0.420-0.533). LLM disagreements were more systematically patterned across the codebook (Cramer's V = 0.554-0.585) than human disagreements (Cramer's V = 0.196). Simulations forecasted agreement ranging from 0.520 to 0.529 when two or more LLMs were paired with one to two human coders. These forecasted agreement levels for hybrid human-LLM workflows fell within the range of human-versus-consensus agreement.
DISCUSSION: D-study simulations support the use of hybrid human-LLM workflows to reach coding consensus for deductive reasoning tasks through mode responses, an alternative consensus-generating mechanism to traditional adjudication discussions. Future work should examine whether these patterns extend across additional deductive coding contexts and model families.