科研速览 · Science Skim继续刷下去 · Keep skimming →
◇ medRxiv2026-08-10· health informatics

A single-patient task exposes a failure of safety alignment in clinical language models

A. Gorenshtein, E. L. Jia, M. Omar, O. R. Brook, M. Ahmed, J. B. Kruskel, Y. Barash, E. Klang

原始摘要(英文原文)· Original abstract
Safety alignment should persist while a language model performs a task. We tested whether a single-patient triage task suppressed a warning about a second patient. Each case centered on Patient 1; Patient 2's urgent problem appeared only in passing. Sixteen models saw each case twice: once as a general assistant and once while producing a triage record for Patient 1. As general assistants, models warned the caller in 87% of cases; under the task, they did so in 21%. Every model showed a significant decrease. Yet under the task, the record still mentioned Patient 2 in 76% of cases and recommended urgent care in 67%. Across 15 open-weight models, repeating the emergency-care instruction raised the warning rate only to 29%; moving the message-to-caller field to the top raised it to 36%. Current safety alignment did not reliably persist under task assignment.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

A single-patient task exposes a failure of safety alignment in clinical language models — 科研速览 Science Skim