科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ PLoS ONE2025-10-28· Benchmarking

Benchmarking large-language-model vision capabilities in oral and maxillofacial anatomy: A cross-sectional study

Nguyễn Việt Anh, Thi Quynh Trang Vuong, Van Hung Nguyen

原始摘要(英文原文)· Original abstract
BACKGROUND: Multimodal large-language models (LLMs) have recently gained the ability to interpret images. However, their accuracy on anatomy tasks remains unclear. METHODS: A cross-sectional, atlas-based benchmark study was conducted in which six publicly accessible chat endpoints, including paired "deep-reasoning" and "low-latency" modes from OpenAI, Microsoft Copilot, and Google Gemini, identified 260 numbered landmarks on 26 high-resolution plates from a classical anatomic atlas. Each image was processed twice per model. Two blinded anatomy lecturers scored responses, including accuracy, run-to-run consistency, and per-label latency, which were compared with χ² and Kruskal-Wallis tests. RESULTS: Overall accuracy differed significantly among models (χ² = 73.2, P < 0.001). OpenAI o3 achieved the highest correctness (53.1%), outperforming its sibling GPT-4o and both Copilot variants, but required the longest inference time. Musculoskeletal structures were recognised more accurately than neurovascular targets, reflecting the greater visual complexity of fine vessels and nerves. Consistency ranged from 43.5% (Gemini Flash) to 65.0% (GPT-4o); deeper modes improved stability for Copilot and Gemini but not accuracy. Median per-label latency spanned three orders of magnitude, from 0.5 s for Gemini Flash to 33 s for o3. CONCLUSIONS: Currently, publicly available multimodal LLMs can only moderately identify oral and maxillofacial landmarks, and no endpoint is sufficiently reliable to serve as a stand-alone answer key. Higher accuracy was achievable with a trade-off in latency, highlighting the need for domain-specific tuning and human oversight. This atlas benchmark study introduced here provides a reproducible yardstick for future model refinement and educational integration.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Benchmarking large-language-model vision capabilities in oral and maxillofacial anatomy: A cross-sectional study — 科研速览 Science Skim