科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Cutaneous and ocular toxicology2026-09-23

Beyond accuracy: an educational benchmarking study of task fragility and reasoning stability of large language models on dermatology board-style questions.

Ahmet Uğur Atılan, Nıyazı Çetın

一句话结论 · In one sentence

Overall, LLM performance exhibited sensitivity to task design, difficulty, and linguistic framing within this educational assessment framework. These findings suggest task fragility and warrant caution in unsupervised educational applications. Practically, this highlights the value of format-aware evaluation when considering LLMs for medical examination preparation or automated question generation, helping to align expectations with their observed educational capabilities.

原始摘要(英文原文)· Original abstract
BACKGROUND: Large language models (LLMs) are increasingly used in medical education and assessment frameworks. However, their reliability under varying task demands remains unclear. Significant gaps exist in the literature regarding the reasoning stability of these models when faced with fluctuations in question structure, difficulty, and linguistic framing. In specialized fields such as dermatology, where diagnostic precision is critical, addressing these inconsistencies is essential for the reliable integration of AI into educational training environments. METHODS: This study was designed as an educational benchmarking analysis under controlled assessment conditions. We evaluated 4 LLMs (ChatGPT 5.0, Claude 4.5 Sonnet, Gemini 2.5 Flash, DeepSeek V3) on 136 dermatology board-style multiple-choice questions, each adapted into 4 task methods: (1) original single-best-answer format; (2) correct option replaced with "None of the answers;" (3) four correct options, requiring identification of the single incorrect option; and (4) no correct option (four distractors). Items were classified as Easy, Medium, or Hard and as Positive or Negative stems. The primary outcome was Correct Response (1/0) across 2,176 model-item-method observations. A generalized linear mixed-effects model with random intercepts for Question ID estimated the effects of model, method, difficulty, and polarity, including interactions; Holm adjustment controlled for multiple comparisons. RESULTS: Overall accuracy was 52%. Using ChatGPT as the reference, all models showed significantly lower odds of correctness (Claude OR, 0.45; Gemini OR, 0.56; DeepSeek OR, 0.54; all p < 0.001). Task method had the largest effect, with marked reductions from Method 1 to Methods 2-4 (Method 4 OR, 0.026; p < 0.001). Negative stems lowered accuracy (OR, 0.57; p < 0.001). Method × Difficulty, Method × Polarity, and Difficulty × Polarity interactions were significant, as was the three-way interaction (p = 0.006), with the steepest losses for medium-difficulty negative items in Methods 2 and 4. CONCLUSION: Overall, LLM performance exhibited sensitivity to task design, difficulty, and linguistic framing within this educational assessment framework. These findings suggest task fragility and warrant caution in unsupervised educational applications. Practically, this highlights the value of format-aware evaluation when considering LLMs for medical examination preparation or automated question generation, helping to align expectations with their observed educational capabilities.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Beyond accuracy: an educational benchmarking study of task fragility and reasoning stability of large language models on dermatology board-style questions. — 科研速览 Science Skim