Bing Li, Jinguo Pang, Yaxin Cai, Jing Yu, Yong Liu, Renrui Tian, Peng Yu, Xiaoling Liang, Shiqing Wang, Peiyuan Wu, Lulu Yang, Lingxiao Wu, Zengxiao Zhao, Li Zheng, Kangning Zhang, Liping Wang, Li Yang, Laurent Itti, Shixian Wen
As a successful proof-of-concept, PA-VLLF demonstrates that vision-language models can successfully internalize expert-derived clinical reasoning. By serving as a standardized, consensus-based intelligent assistant, it reduces the cognitive burden of complex visual evaluations and mitigates individual rater variance. This framework holds promise for improving pain management strategies, supporting highly consistent, AI-assisted clinical decision-making. To support future research, the CPA dataset will be accessible upon the publication of this article to qualified researchers under data privacy regulations.
BACKGROUND: Accurate pediatric pain assessment is essential for effective pain management and procedural safety. However, current evaluations largely rely on subjective scales and casual behavioral observations. Although automated pain assessment methods have been proposed, they often remain complex and less reliable than expert clinicians. This study aimed to establish a proof-of-concept for a novel Pain Assessment Vision-Large Language Framework (PA-VLLF) for consistent pediatric pain evaluation.
METHODS: In this observational proof-of-concept study, we developed and validated the PA-VLLF using representative keyframes extracted from real-world surveillance video recordings of venipuncture from two camera angles. The foundation framework combined ChatGPT-4o and Qwen2-VL vision-language architectures, with customized prompts and fine-tuning to generate FLACC (Facial, Legs, Activity, Cry, Consolability) scores within a human-in-the-loop framework. Data were collected and analyzed from September to December 2024. The study was approved by the institutional review board (IRB no. 294A01) and parental consent was obtained.
RESULTS: We established a Clinical Pain Assessment (CPA) dataset of 1,248 video segments from 104 children, independently annotated by five certified pain experts over a cumulative 950 person-hours to establish a consensus baseline. In this dataset, PA-VLLF-ChatGPT successfully aligned with the collective expert consensus, achieving 86.36% accuracy on expert-labeled pain scores, performing comparably to senior experts (84.69%, p > 0.05) and outperforming machine-learning approaches in our comparisons; precision of pain-level assessment was evaluated using complementary metrics (MAX agreement and Z-Score).
CONCLUSIONS: As a successful proof-of-concept, PA-VLLF demonstrates that vision-language models can successfully internalize expert-derived clinical reasoning. By serving as a standardized, consensus-based intelligent assistant, it reduces the cognitive burden of complex visual evaluations and mitigates individual rater variance. This framework holds promise for improving pain management strategies, supporting highly consistent, AI-assisted clinical decision-making. To support future research, the CPA dataset will be accessible upon the publication of this article to qualified researchers under data privacy regulations.