Diana Vossen, Xenofon Baraliakos, Maria Polyzou
Specialized pregnancy counselling improves maternal and infant outcomes in women with inflammatory rheumatic and musculoskeletal diseases (iRMDs). As such services are not universally available, large language models (LLMs) may increasingly serve as an accessible source of information for patients seeking guidance on family planning, pregnancy, and lactation. This highlights the need to evaluate this software in terms of their adequacy, accuracy and overall reliability. The aim of the analysis is to evaluate the possibility of using LLMs as reliable sources of information regarding female patients suffering from iRMD and going through the period of pregnancy. 54 questions were formulated, which concern the impact and general relationship between pregnancy and rheumatic diseases. The questions were given to four LLMs, ChatGPT, DeepSeek, Gemini and Grok, and the corresponding answers were obtained. The responses were rated across three clinical domains on a 5-point Likert scale (ranging from 1 = strongly disagree/very poor to 5 = strongly agree/excellent) by ten European rheumatology experts and with the help of appropriate statistical indicators calculated, the reliability of the LLMs was assessed. 2160 items were evaluated. All models demonstrated high performance in "Clinical Safety" (Means > 4.10). For "Clinical Accuracy", DeepSeek and ChatGPT scored highest (Median 4.40 for both). Significant deviations were observed in "Completeness and Relevance", where DeepSeek dominated (Mean 4.24, 95% CI 4.17-4.30) and Gemini showed the lowest performance (Mean 3.46, 95% CI 3.27-3.64). Inter-rater agreement was poor to slight (Fleiss' kappa range: -0.043 to 0.059) and moderate for ICC (range: 0.259 to 0.692). LLMs currently exhibit uniformly strong safety behaviors when providing medical information about family planning, pregnancy and lactation in iRMD. However, significant differences exist regarding completeness and relevance, where DeepSeek and ChatGPT are currently the most reliable and complete models for this domain. Furthermore, the variability in responses and low inter-rater agreement highlight that LLMs require further parameterization, are useful tools for routine care, but cannot replace expert medical judgment.