Beyond Simulations: What 20,000 Real Conversations Reveal About Mental Health AI Safety
Current safety evaluations of mental health–focused AI systems predominantly rely on small-scale simulated benchmarks, which inadequately capture the linguistic and contextual diversity of real-world scenarios. This study presents the first systematic safety assessment combining replication across four established benchmarks with an ecological audit of 20,000 real user conversations, comparing specialized mental health AI against six state-of-the-art general-purpose large language models on high-risk topics. Employing clinical expert blind review, LLM-based adjudicators, automated crisis resource triggering, and statistical confidence interval analysis, the findings reveal that the specialized system exhibits significantly lower rates of harmful content in response to prompts involving self-harm, eating disorders, and substance abuse. In live deployment, it achieved zero end-to-end missed detections, with the LLM adjudicator demonstrating 100% sensitivity and 99.2% specificity. The work advocates for ecological auditing as a critical complement to pre-deployment safety testing.