科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Journal of clinical medicine2026-08-02

How Often Do Large Language Models Agree with Each Other-And with the Truth? A Consensus- and Complexity-Stratified Analysis of Data Extraction for Neuroimaging AI.

Nafiye Sanlier, Umid Sulaimanov, Ariorad Moniri, Behman Demir, Gular Ismayilova, Melih Yucel Sanlier, Ugur Erginoglu, Ahmed Rasim Bayramoglu, Maryam Sabah Al-Jebur, Simon Gashaw Ammanuel, Erkin Otles, Abdullah Keles, Ufuk Erginoglu, Mustafa K Baskaya

原始摘要(英文原文)· Original abstract
Background: The reliable integration of large language models (LLMs) into neuroimaging data extraction workflows remains unresolved. Prior benchmarking shows that exact-match accuracy underestimates LLM extraction performance, but whether inter-model consensus and variable complexity can guide automation remains unclear. We evaluated whether inter-model consensus can serve as a confidence signal for human-artificial intelligence (AI) extraction and can guide complexity-stratified workflow triage. Methods: Four frontier LLMs were queried via OpenRouter with an identical zero-shot structured prompt to extract 22 predefined variables from 91 peer-reviewed neuroimaging AI articles, yielding 2002 article-variable items per model. Variables were stratified a priori into low- (n = 7), medium- (n = 8), and high-complexity (n = 7). Performance was compared with an expert reference using exact-match and semantic-equivalence accuracy. Item-level consensus and five triage strategies characterized the efficiency-accuracy trade-off. Results: Semantic-equivalence accuracy converged to 80.5-83.4% across models despite approximately ten percentage-point exact-match differences. Unanimous 4/4 consensus occurred in 45.6% (910/1994) of items, with exact-match accuracy of 85.8%, rising to 95.3% after semantic normalization; however, 14.2% still failed to match the reference. Reliability was complexity-dependent: 96.6% for low-complexity variables, 73.2% for medium-complexity variables, and 38.1% for high-complexity variables. A hybrid strategy auto-accepting 4/4 items and routing 3/4 items to rapid verification reduced estimated review effort by approximately 59%. Conclusions: Inter-model consensus is useful, but it is incomplete and depends on variable complexity. We show that LLM-assisted extraction in neuroimaging AI is a complexity-stratified workflow design problem: low-complexity neuroimaging variables may be selectively automated, while medium-complexity variables require rapid verification, and high-complexity methodological variables should remain human-led.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

How Often Do Large Language Models Agree with Each Other-And with the Truth? A Consensus- and Complexity-Stratified Analysis of Data Extraction for Neuroimaging AI. — 科研速览 Science Skim