科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ JOR spine2026-09-01

Alignment Between Large Language Models and Consensus in Endoscopic Spinal Disc Surgery: A Comparative Analysis Against the Neurocore-SENSED.

Abdullah İyigün, Göker Yurdakul, Alper Dünki, Mehmet Ali Yayla

一句话结论 · In one sentence

Current LLMs reproduce the broad rank ordering of expert endorsement in ESS reporting but cannot reliably distinguish endorsed from contested propositions, principally because of a positive response bias.

原始摘要(英文原文)· Original abstract
BACKGROUND: The Neurocore-SENSED framework, derived from a three-round modified Delphi process involving 77 international spine surgeons, provides a structured reference standard for reporting in endoscopic spine surgery (ESS) for disc disease. The ability of large language models (LLMs) to reproduce graded levels of expert agreement within a reporting framework has not been examined in ESS. OBJECTIVE: To evaluate the extent to which three contemporary LLMs align with the Neurocore-SENSED consensus and whether they discriminate between items of differing consensus level. METHODS: In January 2026, each of the 166 framework items with published item-level endorsement data was submitted once to GPT-5.2, Gemini 3 Pro, and Claude Sonnet 4.6, in independent sessions with web retrieval disabled. Models selected one of five ordered response options directly. Alignment was assessed by Spearman correlation with panel endorsement and by four measures of categorical agreement, each with bootstrap 95% confidence intervals. Separate protocols examined response stability, response format, and prior exposure to the consensus. RESULTS: All three models correlated significantly with panel endorsement. Gemini 3 Pro aligned most closely (r s  = 0.709, 95% CI 0.627-0.776), exceeding Claude Sonnet 4.6 (r s  = 0.622; p = 0.009) and GPT-5.2 (r s  = 0.565; p < 0.001), which did not differ. Categorical agreement was limited and equivalent across models, with observed agreement of 0.506-0.518 and overlapping confidence intervals for all marginal-robust statistics. Between 81.3% and 95.8% of responses fell in the top two categories, and discordance was almost entirely unidirectional: 31 items were endorsed by every model but by fewer than half the panel, against one item in the opposite direction. Weighted kappa was not interpretable in this setting. CONCLUSION: Current LLMs reproduce the broad rank ordering of expert endorsement in ESS reporting but cannot reliably distinguish endorsed from contested propositions, principally because of a positive response bias.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Alignment Between Large Language Models and Consensus in Endoscopic Spinal Disc Surgery: A Comparative Analysis Against the Neurocore-SENSED. — 科研速览 Science Skim