🤖 AI Summary
Current safety evaluations of AI chatbots for children often rely on unvalidated surface-level indicators—such as simplistic refusals—and lack grounding in real-world high-risk scenarios faced by adolescents. This study addresses this gap by conducting in-depth interviews with 19 frontline youth practitioners, including social workers and psychotherapists, to systematically incorporate professional practitioner judgment into AI safety assessment for the first time. Centered on authentic harm scenarios and adolescent needs, the research establishes a novel evaluation paradigm. Through qualitative and contextualized analysis of AI responses across representative high-risk situations, the study identifies behavioral patterns prone to causing harm as well as characteristics of supportive responses. These insights inform concrete recommendations for refining AI safety evaluation frameworks and response mechanisms.
📝 Abstract
Youth increasingly turn to AI chatbots for social and emotional support, raising concerns about how these systems respond, especially in high-stakes situations. However, existing child safety evaluations of AI lack grounding in real-world harms that youth experience, rely on unvalidated assumptions about what counts as an appropriate output (e.g., refusal), and typically focus on detecting adversarial prompts or surface-level harms in outputs only. Thus, these evaluations can fail to detect responses that pose harm to youth in practice. To better understand the limitations of current evaluation practices, we conducted interviews with 19 practitioners working directly with youth in vulnerable situations, including social workers, therapists, and psychologists, asking them to reflect on chatbots' responses to risky situations commonly faced by youth, as established in prior empirical work. Practitioners identified chatbot behaviors likely to cause harm as well as those that could meaningfully support youth in difficult moments, discussed the role that chatbots should (and should not) play in these interactions, and offered concrete recommendations for improving chatbot responses. Based on these findings, we provide recommendations for AI child safety evaluation and infrastructure, and highlight the need for incorporating practitioners' perspectives into safety work.