Zhiwei Lin, Xiaoying Huang, Yanhong Yan, Xinglin Zheng, Wenyu Wang, Guang Su
This evidence-based review supports greater consistency in reporting and organizing LLM applications within nursing research and practice.
AIM: To systematically identify and categorize the outcomes, evaluation metrics, and measurement tools employed in studies evaluating the application of large language models (LLMs) in the nursing field.
BACKGROUND: LLMs are increasingly applied in nursing education, clinical practice, and research. However, outcomes, evaluation metrics, and measurement tools used to evaluate these applications remain heterogeneous and inconsistently reported, limiting cross-study comparability and evidence synthesis.
METHODS: This scoping review was conducted in accordance with the Joanna Briggs Institute methodology and PRISMA-ScR guidelines. Outcomes, evaluation metrics, and measurement tools were extracted and categorized according to the Consensus-based Method for the Selection of Health Outcome Measurement Tools (COMET) framework domains, then mapped to LLM task types based on the Trustworthiness and Reporting of Items for Prediction Model Development-Large Language Models (TRIPOD-LLM) guidelines.
SOURCES OF EVIDENCE: Eight English and Chinese databases were searched to identify original empirical studies evaluating LLM applications in nursing, yielding 50 studies that met the inclusion criteria.
RESULTS: A total of 56 distinct evaluative elements (encompassing clinical/educational outcomes, technical evaluation metrics, and standardized measurement tools) were extracted and categorized into three COMET-derived domains: objective efficiency indicators, intermediate effect indicators, and subjective experience indicators. Most measures focused on text generation and summarization tasks, while evaluation metrics for data extraction and decision support tasks were relatively limited. This reflects issues of reporting inconsistency and insufficient standardization.
DISCUSSION: Performance metrics for LLMs in nursing applications exhibit uneven distribution across task modalities and operational scenarios, underscoring the necessity for standardized evaluation guidelines.
CONCLUSION: This evidence-based review supports greater consistency in reporting and organizing LLM applications within nursing research and practice.
IMPLICATIONS FOR NURSING PRACTICE: The findings of this review can support nurses and educators in selecting appropriate evaluation indicators, promoting safe and evidence-informed integration of LLMs.
IMPLICATIONS FOR NURSING POLICY: When evaluating LLM applications in nursing, enhancing the clarity and consistency of outcome metric reporting facilitates the development of institutional-level guidelines.