Matthew Flathers, Phuong Anh Nguyen, Julian Herpertz, Mason Granof, Sean Ryan, Leah Wentworth, Christine Yu Moutier, John Torous
A single benchmark score can support markedly different claims depending on the assumed standard of clinical behavior, the instrument's remaining measurement range, and the configuration that produced the result. The skills required to make these distinctions must become core competencies. Benchmark results are increasingly utilized to support claims about mental health safety that may not be accurate, making it necessary to close the gap between clinical measurement and AI.
BACKGROUND: Millions of people use language models to discuss mental health concerns, including suicidal ideation, but limited frameworks exist for evaluating whether these systems respond safely. Benchmarking, the practice of administering standardized assessments to language models, offers direct parallels to clinical competency evaluation, yet few clinicians are involved in designing, validating, or interpreting these assessments.
AIMS: To introduce mental health professionals to benchmarking language models by administering a validated clinical instrument and demonstrating how configuration decisions, measurement limitations, and scoring context affect result interpretation.
METHOD: We administered the Suicide Intervention Response Inventory (SIRI-2) programmatically to nine commercially available language models from three providers. Each item was presented 60 times per model (three prompt variants × two temperature settings x 10 repetitions), yielding 27,000 individually scored responses compared against point-in-time expert consensus.
RESULTS: Total scores ranged from 19.5 to 84.0 (expert panel baseline: 32.5). Prompt design alone shifted individual model scores by as much as the difference between trained and untrained human groups. The best performing model approached the instrument's measurement floor. All nine models consistently overrated clinically inappropriate responses that sounded supportive.
CONCLUSIONS: A single benchmark score can support markedly different claims depending on the assumed standard of clinical behavior, the instrument's remaining measurement range, and the configuration that produced the result. The skills required to make these distinctions must become core competencies. Benchmark results are increasingly utilized to support claims about mental health safety that may not be accurate, making it necessary to close the gap between clinical measurement and AI.