科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Substance use & misuse2026-08-19

Evaluating Large Language Models in Response to Questions on Substance Use: Helpful or Harmful?

Samuel Maddams, Shan Chen, Danielle S Bitterman, Samata Sharma, David B Hathaway, Patrick L McGuire, Ian G Johnson, Alexander C Wu, Sara Prostko, Joji Suzuki

一句话结论 · In one sentence

While many LLMs provided competent and safe responses, none were completely competent and non-stigmatizing, highlighting the potential but also ongoing need for refinement and expert verification.

原始摘要(英文原文)· Original abstract
BACKGROUND: Individuals with substance use disorders (SUD) are obtaining health-related information from various large language models (LLMs). We aimed to assess whether LLMs provide responses concordant with the current evidence base and whether they provide harmful responses. METHODS: Twenty questions related to SUD were posed to three LLMs (Gemini-1.5-pro-001, Claude-3-5-sonnet, and GPT-4) in May 2024. Each response was independently rated by three experienced addiction specialists, and disagreements were resolved by two additional experienced addiction specialists. All raters were blinded to the LLM. Each rater assessed whether (I) a competent addiction specialist would agree with the response, (II) the response contained stigmatizing language as defined by National Institute on Drug Abuse, or (III) the response contained harmful content. RESULTS: 88% of responses were rated as competent and 92% as not harmful. Gemini-1.5-pro-001 had the highest rate of competence (95%), followed by Claude-3-5-sonnet and GPT-4 (both 85%). Gemini-1.5-pro-001 produced no harmful responses, while Claude-3-5-sonnet and GPT-4 produced 10% and 15%, respectively. 30% of responses from both Gemini-1.5-pro-001 and Claude-3-5-sonnet had contained stigmatizing language, compared to 10% for GPT-4. CONCLUSIONS: While many LLMs provided competent and safe responses, none were completely competent and non-stigmatizing, highlighting the potential but also ongoing need for refinement and expert verification.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Evaluating Large Language Models in Response to Questions on Substance Use: Helpful or Harmful? — 科研速览 Science Skim