Alessandro Polizzi, Luigi Nibali, Pasquale Santamaria, Pietro Venezia, Gianluca Tartaglia, Gaetano Isola
Benchmarking AI chatbots using structured questions derived from the EFP S3-level guideline for stage IV periodontitis showed high clarity but variable guideline alignment and source reporting. Apparent chatbot rankings were attenuated after excluding Sources, and chatbot identity was not associated with total QAMAI score after adjustment. Rankings may therefore be influenced by retrieval and source-reporting features rather than reflecting intrinsic model capability. All outputs require clinician verification.