Chang-Ki Min
Accuracy was strongly affected by model choice, information type and user-suggested diagnoses, but no model reached a level that would be acceptable for clinical use.
OBJECTIVES: To compare the diagnostic performance of three multimodal AI chatbots on oral and maxillofacial radiographic images and to examine how additional information, delivery mode, and user-suggested diagnoses affect accuracy.
METHODS: Three AI chatbots (GPT-5.1, Gemini 3 Flash, Claude Opus 4.7) were tested on 90 cases comprising normal controls, osteomyelitis, and benign jaw lesions. Inputs were combinations of a panoramic image, a cropped panoramic image, an axial CBCT image or a text-based cue. Inputs were delivered all at once or sequentially. Accuracy was scored at category and specific-diagnosis levels using non-parametric tests with false-discovery-rate correction.
RESULTS: With the panoramic image alone, accuracy for diseased cases was low (0-60%) but rose to as high as 38-92% in each model's best condition with added information. A cropped image was the most consistently beneficial additional visual input, whereas an axial CBCT image provided less improvement. GPT-5.1 recognised normal cases well but missed lesions, Gemini 3 was sensitive but less specific, and Claude 4.7 defaulted to benign diagnoses. Correct verbal cues increased accuracy, whereas a misleading cue caused decline. Gemini 3 accepted a false benign suggestion in 96% of cases it had initially classified as normal. For benign lesions, specific-diagnosis accuracy was almost half of category-level accuracy.
CONCLUSIONS: Accuracy was strongly affected by model choice, information type and user-suggested diagnoses, but no model reached a level that would be acceptable for clinical use.
ADVANCES IN KNOWLEDGE: This multi-model comparison isolates the effects of information type, delivery mode, and sycophancy in oral and maxillofacial radiology.