科研速览 · Science Skim继续刷下去 · Keep skimming →
◇ medRxiv2026-08-23· health informatics

MISP-Bench: Decomposing User-Provided False Priors into Answer, Rationale, and Guard Effects

I. Jeong, Y. Kim, J.-H. Park, H. Lee

原始摘要(英文原文)· Original abstract
Large language models often agree with a user's wrong belief. Existing benchmarks measure how often this happens, but not which part of the belief causes the harm: the wrong answer the user states, the reason they give for it, or only the two together. We introduce MISP-Bench, which decomposes a user-provided false prior into these parts. Holding the question and the correct answer fixed, we vary each part on its own across 1,724 audited items (1,430 medical multiple-choice, 294 free-form math), 10 open-weight models (1B-27B), and 13 prompt conditions. Three findings stand out. First, the wrong answer and the wrong rationale together do less damage than the sum of each alone (sub-additive on 7 of 10 models), so a defense can target either part. Second, the popular safety prompt "verify the reasoning first" does not help uniformly: it splits models into recovery, no-effect, and reversal groups, and makes outcomes worse on 4 of 10; a within-model reasoning-budget sweep suggests the split is driven by reasoning capacity. Third, two attacks with the same accuracy drop can involve very different behavior: models adopt the user's wrong answer 78% of the time when the distractor matches an error-prone option identified by a strong reference model, but only 39% when it does not, so reporting accuracy loss alone misses how the model fails. Model size does not account for any of these in our 1B-27B range. The findings replicate on two frontier models (GPT-5.6, Gemini-2.5-Pro) and survive regenerating the distractors with a different model. We release the corpus, all response records, and a reusable six-category audit that removed 770 flawed source items (732 of them multi-answer items that single-best-answer scoring cannot handle). Data: https://huggingface.co/datasets/yh0502/misp-bench (CC-BY-4.0); code: https://github.com/anon-misp-2026/misp-bench.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

MISP-Bench: Decomposing User-Provided False Priors into Answer, Rationale, and Guard Effects — 科研速览 Science Skim