K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations
为评估大型语言模型在高风险心理健康对话中的安全性,开发了K-Bench基准测试,通过临床校准和多情景模拟来评价模型表现。
为评估大型语言模型在高风险心理健康对话中的安全性,开发了K-Bench基准测试,通过临床校准和多情景模拟来评价模型表现。
This study investigates the trade-off between clinical safety and environmental impact in therapeutic large language models. By integrating K-Bench clinical safety scores with EcoLogits life cycle assessment, it performs a fine-grained analysis of 47 model configurations across four environmental dimensions: energy consumption, carbon emissions, water use, and abiotic resource depletion. The work reveals, for the first time, that within high-safety regimes, marginal gains in clinical safety incur nonlinear surges in environmental costs—specifically, a mere 2.61-point increase in safety score corresponds to approximately a 60-fold rise in energy consumption, with additional inference compute not necessarily yielding further safety benefits. To address this, the study proposes dynamic model selection strategies, such as model cascading, which can maintain performance in high-risk clinical scenarios while substantially reducing environmental footprint.
为评估大型语言模型在高风险心理健康对话中的安全性,开发了K-Bench基准测试,通过临床校准和多情景模拟来评价模型表现。
This study investigates the trade-off between clinical safety and environmental impact in therapeutic large language models. By integrating K-Bench clinical safety scores with EcoLogits life cycle assessment, it performs a fine-grained analysis of 47 model configurations across four environmental dimensions: energy consumption, carbon emissions, water use, and abiotic resource depletion. The work reveals, for the first time, that within high-safety regimes, marginal gains in clinical safety incur nonlinear surges in environmental costs—specifically, a mere 2.61-point increase in safety score corresponds to approximately a 60-fold rise in energy consumption, with additional inference compute not necessarily yielding further safety benefits. To address this, the study proposes dynamic model selection strategies, such as model cascading, which can maintain performance in high-risk clinical scenarios while substantially reducing environmental footprint.