科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Cornea2026-08-04

Reproducibility of Tear Ferning Test Classification by Human Examiners and Artificial Intelligence Models: A Comparative Study.

Sérgio Felberg, Guilherme D'Agosto Bernardes, Felipe Augusto Casseb Dos Santos, Lucas Della Paolera, Elcio Hideo Sato, Bernardo Kaplan Moscovici, Paulo Dantas

一句话结论 · In one sentence

Binary classification was more reproducible than the 4-grade Rolando scale for humans and platforms, but it is a pragmatic standardization strategy, not a replacement for ordinal grading. Claude best matched the human benchmark, whereas ChatGPT and Gemini were less reproducible and systematically underclassified abnormalities. Platform assistance may have adjunctive value, but clinical use will require external validation, accuracy and repeated-session testing, and safety evaluation.

原始摘要(英文原文)· Original abstract
PURPOSE: To compare the reproducibility of tear ferning classification by human evaluators and 3 multimodal artificial intelligence platforms using the 4-grade Rolando scale and a predefined binary scheme, and to clarify its clinical role. METHODS: Eighty polarized-light microscopy tear-film images were independently graded (Rolando grades I-IV) by 3 ocular surface researchers and 3 platforms (ChatGPT, Claude, Gemini), each evaluated with an identical prompt and image order in single-run conditions. Grades I-II were defined a priori as normal and III-IV as abnormal. RESULTS: For the 4-grade scale, Fleiss kappa was 0.65 (humans), 0.44 (6 raters), and 0.37 (platforms); binary reclassification raised these to 0.76, 0.66, and 0.66. Against consensus, Claude agreed best (binary agreement 93.8%, kappa 0.87; 4-grade 82.5%, weighted kappa 0.84, 95% confidence interval, 0.75-0.91), within the interhuman range. ChatGPT and Gemini showed lower 4-grade agreement (weighted kappa 0.48 each) and significantly underclassified abnormalities (McNemar P = 0.0004 and 0.00006), whereas Claude showed no such asymmetry (P = 0.3750). CONCLUSIONS: Binary classification was more reproducible than the 4-grade Rolando scale for humans and platforms, but it is a pragmatic standardization strategy, not a replacement for ordinal grading. Claude best matched the human benchmark, whereas ChatGPT and Gemini were less reproducible and systematically underclassified abnormalities. Platform assistance may have adjunctive value, but clinical use will require external validation, accuracy and repeated-session testing, and safety evaluation.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Reproducibility of Tear Ferning Test Classification by Human Examiners and Artificial Intelligence Models: A Comparative Study. — 科研速览 Science Skim