Assessing AI Companion Chatbot Safety for Vulnerable Users: A Crisis-Stage Multi-Platform Evaluation
Abstract
AI companion chatbots are increasingly used by vulnerable individuals, including adolescents in suicidal crisis, as an emotional support channel. However, there are no systematic framework to compare their safety behavior regarding self-harm and suicide risks. These risks are not hypothetical: several high-profile cases have linked companion chatbots to adolescent suicide deaths, prompting regulators in multiple countries to assess their impact on vulnerable users. Existing safety evaluations focus on base model properties leaving three gaps: product-layer behaviors shaped by engagement-optimizing design are unmeasured; the ideation-to-planning transition as the most clinically consequential escalation point in suicidal crisis, has no dedicated framework; and companion apps have not been benchmarked alongside general-purpose LLMs under identical conditions. We introduce MIRROR-Eval (Multi-turn Interaction Risk Review: Observing suicidal-crisis Responses) with a three-dimension rubric grounded in psychological studies and clinical practices, including covering user risk severity, chatbot risk awareness, and intervention behavior evaluated using a validated LLM-as-a-Judge pipeline with expert annotators. We observe a three-tier harm structure: companion apps (Character.AI and Replika) and Grok 4.3 form a high-harm tier; Gemini 3.1 Pro is mid-range; GPT-5.5 and Claude Sonnet 5 form a clearly safer tier whose residual any-flag rates reflect isolated miscalibrations rather than systematic harm. Companion harm is structural and crisis-agnostic: companion apps average 15.1 distinct harmful behaviors per trial versus 5.0 for general LLMs. At the planning stage, companion apps produce an escalation behaviors combination (glorification and normalisation, advancing toward harmful action, and harmful methods and instructions), near-zero at ideation. Dataset and code are released at [https://huggingface.co/datasets/imda-biztech/MIRROR-Eval].