科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ The Journal of prosthetic dentistry2026-09-08

Comparative evaluation of large language models in interpreting the scientific literature on intraoral scanners across varying input levels.

Pranay Jain, Fathima Banu R, Anand Kumar V, Sarvesh Kulkarni, Maheshwaran Ks

一句话结论 · In one sentence

The Gemini-3.0 Flash responses excelled in highly structured, direct-response questions to relevance, clarity, and closeness to the article, whereas the ChatGPT-5.2 response demonstrated greater scientific accuracy at all levels and performed better with simplified prompts that require high logical reasoning. Framing appropriate prompt input is essential for achieving better scientific accuracy when using AI tools for research analysis.

原始摘要(英文原文)· Original abstract
STATEMENT OF PROBLEM: Dental professionals increasingly seek artificial intelligence (AI)-generated sources for concise information on emerging topics in advancement and scientific literature; however, concerns persist regarding the accuracy of AI tools and the reliability of responses generated based on the complexity of the prompt. PURPOSE: The purpose of this study was to evaluate and compare the response accuracy of 2 advanced large language models (LLMs): OpenAI ChatGPT-5.2 (Auto) and Google Gemini-3.0 Flash in interpreting and analyzing published data on intraoral scanners. MATERIAL AND METHODS: A total of 20 articles were selected through a structured PubMed search. Two independent senior prosthodontists were blinded and assessed every response generated by the LLMs at each prompt level. Prompts were fed at different levels (1-4) of complexity: Level 1 consisted of simplified prompts requiring high logical reasoning by the LLM, while Level 4 consisted of a structured prompt including methodology. The responses were evaluated across the 4 domains of relevance, scientific accuracy, clarity, and closeness to article content using a 5-point Likert scale. The nonparametric Friedman test and Wilcoxon signed-rank test with Bonferroni correction were conducted for statistical analysis (α=.05). RESULTS: The median output for ChatGPT-5.2 for relevance of answer and clarity were higher than for Gemini-3.0 at the simpler prompt levels of Level 1 and Level 2 (P<.05), whereas the trend reversed at higher levels, with Gemini-3.0 performing significantly better at Level 3 (P<.001) and Level 4 (P<.001). In terms of scientific accuracy, ChatGPT-5.2 consistently achieved higher scores at Level 1 (P=.003) and Level 4 (P=.023), whereas Gemini-3.0 plateaued across levels (P>.05). In contrast, Gemini-3.0 outperformed ChatGPT-5.2 in closeness to article at Level 2 (P<.001) and 4 (P=.019). CONCLUSIONS: The Gemini-3.0 Flash responses excelled in highly structured, direct-response questions to relevance, clarity, and closeness to the article, whereas the ChatGPT-5.2 response demonstrated greater scientific accuracy at all levels and performed better with simplified prompts that require high logical reasoning. Framing appropriate prompt input is essential for achieving better scientific accuracy when using AI tools for research analysis.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Comparative evaluation of large language models in interpreting the scientific literature on intraoral scanners across varying input levels. — 科研速览 Science Skim