科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ La Tunisie medicale2026-05-01

Comparative benchmarking of three artificial intelligence chatbots (ChatGPT-5.1, Qwen- 3 Max, and Perplexity AI) on faculty validated undergraduate respiratory physiology multiple-choice examinations.

Khouloud Kchaou, Salma Mokaddem, Soumaya Khaldi, Yacine Ouahchi, Khadija Ayed, Rym Baati

一句话结论 · In one sentence

Although all evaluated AI chatbots performed well on respiratory physiology MCQs, substantial variability was observed across models. These findings indicate that AI chatbots are not interchangeable and underscore the need for institution-level evaluation using local examination material before their integration into medical education.

原始摘要(英文原文)· Original abstract
BACKGROUND: Artificial intelligence (AI) chatbots are increasingly used by medical students for learning and examination preparation. However, their reliability in mechanistically demanding disciplines such as respiratory physiology remains insufficiently evaluated using authentic examination material. This study aimed to assess and compare the performance of three AI chatbots on validated undergraduate respiratory physiology examination questions. METHODS: We conducted a cross-sectional analytical study including all respiratory physiology multiple-choice questions (MCQs), each comprising five propositions, used in official undergraduate examinations at the Faculty of Medicine of Tunis during the academic years 2023-2024 and 2024-2025 (101 questions). Three AI chatbots (ChatGPT-5.1, Qwen-3 Max, and Perplexity AI) were evaluated using standardized French-language prompts. Exact question-level concordance with the faculty-validated answer key, established by the teaching staff responsible for respiratory physiology examinations, was compared using Cochran's Q test and pairwise McNemar tests. At the proposition level, responses were analyzed as binary outcomes (true/false) to compute accuracy, sensitivity, and specificity. RESULTS: Question-level exact concordance differed significantly across models (Cochran's Q = 10.18, p = 0.006). ChatGPT showed the highest concordance rate (79.2%), followed by Qwen (74.3%) and Perplexity (62.4%). At the proposition level (505 propositions), ChatGPT achieved the highest overall accuracy (94.1%) and sensitivity for true propositions (96.5%), whereas Qwen demonstrated the highest specificity for false propositions (94.2%). CONCLUSION: Although all evaluated AI chatbots performed well on respiratory physiology MCQs, substantial variability was observed across models. These findings indicate that AI chatbots are not interchangeable and underscore the need for institution-level evaluation using local examination material before their integration into medical education.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Comparative benchmarking of three artificial intelligence chatbots (ChatGPT-5.1, Qwen- 3 Max, and Perplexity AI) on faculty validated undergraduate respiratory physiology multiple-choice examinations. — 科研速览 Science Skim