Kento Masukawa, Haruse Oikawa, Lei Dong, Hideyuki Hirayama, Mitsunori Miyashita
BackgroundPatient-reported outcomes (PROs) are crucial for accurate patient assessments. In palliative care, the Integrated Palliative care Outcome Scale (IPOS) is a widely used measure of these outcomes. Nevertheless, manually entering PRO data into electronic medical records or databases can be burdensome. Thus, a previous study developed models to automatically extract IPOS scores from clinical conversation data using speech recognition and traditional natural language processing. However, further improvements in performance are required. This study examined the accuracy of extracting IPOS scores from clinical conversation data using the latest large language models (LLMs).MethodsWe collected conversation data related to IPOS from 100 patients receiving home care at a clinic in Japan between February and May 2023. Two LLMs, Gemini and Copilot, were used to extract IPOS scores. Twenty percent of the data were used to optimize the input instructions for the LLMs, and the remaining 80% were used for performance evaluation. The primary performance metric was the macro-F1 score.ResultsThe mean macro-F1 scores ± standard deviation (SD) for Gemini were 0.96 ± 0.08 for physical symptoms, 0.91 ± 0.11 for emotional symptoms, and 0.84 ± 0.05 for communication/practical issues. For Copilot, the corresponding scores were 0.94 ± 0.08, 0.86 ± 0.10, and 0.85 ± 0.12, respectively.ConclusionThese findings suggest the technical feasibility of LLMs for extracting IPOS scores from clinical conversation data, although further studies are needed to confirm this finding. Integrating speech recognition and LLMs may have significant potential to reduce the burden of sharing PRO-related information.