科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ European heart journal. Digital health2026-08-01

Comparative evaluation of artificial intelligence-assisted literature search tools for identifying clinically meaningful evidence in cardiology.

Riccardo Di Febo, Maximiliano Jeanneret Medina, Alexandre Renaud, Maryam Dridi, Valentine Pecriaux, Benoit Lequeux, Stephane Lafitte, Baptiste Maille, Aymeric Menet

一句话结论 · In one sentence

AI-assisted tools showed heterogeneous performance, ChatGPT-5 performing best in this cardiology setting. These preliminary, context-specific findings support hybrid human-AI strategies in which AI complements rather than replaces transparent database searches such as PubMed; larger-scale, multi-domain studies are needed to confirm and generalize them.

原始摘要(英文原文)· Original abstract
AIMS: The rapid expansion of biomedical literature challenges clinicians' and researchers' ability to identify clinically meaningful evidence. We systematically compared five literature search tools, four artificial intelligence (AI)-assisted and one conventional, across clinically relevant cardiology research scenarios, using a blinded expert-validated gold standard to assess their ability to retrieve relevant and key references. METHODS AND RESULTS: We evaluated ChatGPT-5, Elicit, Consensus, Scite, and PubMed across four cardiology topics defined by maturity and specificity, with multiple standardized prompts. Three electrophysiology experts independently and blindly rated all retrieved references, defining two gold standards: expert-rated relevance and expert-selected key references. ChatGPT-5 achieved the highest proportion of relevant articles (90% [88-100], P < 0.001) and the highest key-reference overlap (60% [43-68], P < 0.001), whereas Scite performed lowest (20% and 10%, respectively). The tool was the primary determinant of performance (partial R 2 = 0.50), whereas prompt formulation had no significant effect. In a pre-specified subanalysis restricted to clinical studies, ChatGPT-5 and human-conducted systematic reviews overlapped by 42% (96% of shared articles highly relevant), with 58% distinct references, indicating complementary AI and human retrieval; ChatGPT-5 produced hallucinated citations when long reference lists were requested for emerging topics, underscoring the need for human verification. CONCLUSION: AI-assisted tools showed heterogeneous performance, ChatGPT-5 performing best in this cardiology setting. These preliminary, context-specific findings support hybrid human-AI strategies in which AI complements rather than replaces transparent database searches such as PubMed; larger-scale, multi-domain studies are needed to confirm and generalize them.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Comparative evaluation of artificial intelligence-assisted literature search tools for identifying clinically meaningful evidence in cardiology. — 科研速览 Science Skim