Lin Chen, Xiuli Tang, Hongbiao Gao, Li Sun, Mi Han, Jing Liang, Yue Xu, Meng Zhang, Meiqing Chen, Ran Cheng, Xiaohui Xiang, Qian Zhao, Yuehe Ma, Ling Liu, Jia Xu
Patient satisfaction and experience surveys are essential for assessing the quality of emergency nursing care, yet conventional questionnaire development is often slow because it relies on multiple rounds of expert consensus, translation, linguistic editing, and pilot testing. To explore whether large language models can streamline this process, we evaluated the feasibility of integrating ChatGPT into the design and preliminary validation of an emergency-department (ED) patient satisfaction and experience questionnaire. A three-phase methodological feasibility study was conducted in 2 tertiary EDs between March and August 2025. In Phase I, ChatGPT-4 and ChatGPT-5 (OpenAI, San Francisco, CA) generated 40 candidate items across 5 service domains. In Phase II, a 10-member Delphi panel rated relevance and clarity and we calculated item- and scale-level content validity indices (I-CVI and S-CVI). In Phase III, 200 ED patients completed both the ChatGPT-generated questionnaire and a nurse-developed comparison tool; reliability (Cronbach's α), item-total correlations, readability, and completion time were compared. ChatGPT produced well-structured items covering waiting time, communication, empathy, environment, professionalism, and perceived safety. The ChatGPT questionnaire achieved an S-CVI of 0.91 (vs. 0.93), with higher internal consistency (α=0.92 vs. 0.89), better readability (Flesch-Kincaid grade 6.8 vs. 8.1; P<.01), and a 32% shorter completion time. EFA and CFA further supported stronger structural validity for the AI-generated instrument than for the nurse-developed comparator. These findings support ChatGPT, particularly GPT-5.0, as a practical aid for rapid, reliable early-stage instrument development in nursing research.