Hsi-Sheng Wei, Prabhleen Atwal, Lillie Humphrey, Yi-An Yu
Background Large language models (LLMs) are increasingly used in mental health contexts, yet their ability to respond safely to suicide-related risk in dynamic interactions remains unclear, and clinically grounded evaluation frameworks are limited. Methods We developed a clinically informed framework based on crisis intervention principles, encompassing connection and listening, suicide risk assessment, and coping strategy exploration. Using a simulation-based design, an AI agent enacted a help-seeking individual across four scenarios with explicit and culturally embedded suicide risk indicators. Three widely used LLMs were evaluated in repeated multi-turn interactions. Responses were coded using binary criteria by three trained raters (intraclass correlation coefficient = 0.78). Results All models demonstrated strong relational engagement, including empathy and active listening. However, suicide risk assessment was inconsistent and often inadequate. Earlier models rarely initiated direct inquiry about suicidal ideation, while newer models showed improvement but lacked consistency. Detection of culturally embedded risk cues remained limited across all models. Coping-related responses were common but generally unstructured and lacked key elements of evidence-based safety planning. Performance was stronger for explicit than for indirect or culturally embedded risk signals. Conclusions LLMs can generate empathetic responses but show critical deficiencies in systematic risk assessment and structured intervention. The proposed framework offers a clinically grounded, transferable approach for evaluating AI safety in high-risk mental health contexts and informing responsible development and governance.