科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Digital Health2026-02-01· Psychology

Exploring evaluation measures of large language models for family caregiver use: A scoping review

Han Soojeong, Hannah Cho, Yong K Choi, Gregory L. Alexander

原始摘要(英文原文)· Original abstract
Background: Large language models have a huge positive impact on various disciplines, including healthcare. As family caregivers are an essential part of the healthcare system, they need support and can benefit from the technology. However, there is no consensus on reliable and valid measures to evaluate large language models. Objective: This study aims to review the literature on the evaluation measures of large language models for caregivers. Methods: We conducted a scoping review guided by Arksey and O'Malley methodology and the PRISMA-ScR checklist. A literature search on PubMed, EMBASE, CINAHL, and PsycINFO, from 2018 through July 2024, was carried out. An additional rapid review was conducted for the recent literature update from July 2024 through November 2025. Results: All 10 final publications that met the inclusion criteria out of 1812 focused on ChatGPT, whereas three of them also addressed other large language models, such as Google Bard and Bing AI. The most commonly assessed core conceptual components of evaluation measures were accuracy, reliability, readability, and comprehensiveness. Overall, the included studies reported that large language models' responses were somewhat accurate and reliable and mixed results in readability and comprehensiveness. The final 14 publications from a rapid review offered additional evidence on ChatGPT-centrism. Conclusions: This review provides a comprehensive overview of the measures for evaluating large language models and highlights the need for their improvement using reliable and valid measures. The findings guide the direction of future research and practice to maximize the benefits through continuous quality improvement.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Exploring evaluation measures of large language models for family caregiver use: A scoping review — 科研速览 Science Skim