M. Uzcategui Salazar, May Khin Chaw, Yvette Hellier, Stephanie L. Hsia, Katherine Gruenberg
OBJECTIVE: Qualitative research remains underutilized in health professions education, in part due to insufficient training and time-intensive analytic methods. Recent advances in generative artificial intelligence offer new opportunities to streamline the qualitative research process using large language models such as GPT-4. However, the accuracy of GPT-4-generated codes and themes remains underexplored in health professions education research. This study characterizes qualitative analyses assisted by a general-purpose GPT-4 compared to traditional human-conducted analyses. METHODS: Two health professions datasets were previously analyzed using content or thematic analysis and then reanalyzed using a version of GPT-4. Researchers compared the accuracy, alignment, relevance, and appropriateness of codebooks and themes produced by GPT-4 with the prior findings. Dichotomous numerical ratings and explanations were assessed independently and then discussed collaboratively to identify strengths and weaknesses associated with GPT-4 qualitative analysis. RESULTS: In total, 36 survey responses and seven 1-h interview transcripts were analyzed using GPT-4. The codebooks and themes generated by GPT-4 generally aligned with human-identified concepts. Challenges included failure to detect low-frequency codes, difficulty constructing coherent code relationships, and a lack of nuance in theme descriptions and quote selection. CONCLUSION: GPT-4 can support, though not replace, human-led qualitative analysis. A general understanding of qualitative research processes and the dataset is necessary for researchers to identify potential gaps, limitations, and redundancies in qualitative findings generated by GPT-4.