科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Journal of educational evaluation for health professions2026-01-01

Refusal-aware evaluation of frontier AI models available in June 2026 using the Japanese National License Examination for Pharmacists: a comparative study.

Hiroyasu Sato, Katsuhiko Ogasawara, Hidehiko Sakurai

一句话结论 · In one sentence

Near-saturation benchmark performance coexisted with frequent and partly run-dependent refusals in a safeguard-equipped model. Overall accuracy, accuracy excluding refusals, refusal rate, and refusal consistency describe complementary aspects of performance; therefore, repeated, refusal-aware evaluation is needed to interpret frontier AI models in pharmacy education.

原始摘要(英文原文)· Original abstract
PURPOSE: Conventional single-run accuracy may be insufficient for evaluating frontier generative artificial intelligence (AI) models when safety-related refusals occur. This 4-model benchmark examined the need for repeated, refusal-aware evaluation using the Japanese National License Examination for Pharmacists (JNLEP). METHODS: ChatGPT GPT-5.5, Gemini 3.5 Flash, Claude Opus 4.8, and Claude Fable 5 were evaluated using all 345 questions from the 107th JNLEP. The original Japanese questions, including image-containing items, were submitted through application programming interfaces (APIs) in 3 independent runs. Refusals were treated as incorrect when overall accuracy was calculated. For Fable 5, accuracy excluding refusals, refusal rate, refusal consistency across runs, the subject-wise distribution of refusals, and system-assigned refusal categories were also evaluated. RESULTS: The mean overall accuracies were 98.7% for GPT-5.5, 98.3% for Gemini 3.5 Flash, 96.1% for Claude Opus 4.8, and 70.3% for Claude Fable 5. Fable 5 had a mean refusal rate of 29.0%, whereas its mean accuracy excluding refusals was 99.0%. Among the 345 items, 93 were refused in all 3 runs, 14 were refused inconsistently across runs, and 238 were never refused. All 300 refusal responses were assigned to the bio category. Refusals were most frequent in Biology (90.0%) and Pharmacology (69.2%) but uncommon in Practice (2.5%). CONCLUSION: Near-saturation benchmark performance coexisted with frequent and partly run-dependent refusals in a safeguard-equipped model. Overall accuracy, accuracy excluding refusals, refusal rate, and refusal consistency describe complementary aspects of performance; therefore, repeated, refusal-aware evaluation is needed to interpret frontier AI models in pharmacy education.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Refusal-aware evaluation of frontier AI models available in June 2026 using the Japanese National License Examination for Pharmacists: a comparative study. — 科研速览 Science Skim