K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations
为评估大型语言模型在高风险心理健康对话中的安全性,开发了K-Bench基准测试,通过临床校准和多情景模拟来评价模型表现。
为评估大型语言模型在高风险心理健康对话中的安全性,开发了K-Bench基准测试,通过临床校准和多情景模拟来评价模型表现。
本文提出了一种基于双阶段MLP-QRNN层次结构的分位数引导特征提取框架,以解决工业制造系统中多时间尺度预测性维护问题。
该论文提出TQRNN30d框架,通过结合量化回归神经网络与多流时间融合分类器,解决长期工业预测维护中的设备退化识别问题。
论文针对LLM在社交机器人中使用时出现的不一致人格问题,通过提出包含八个功能组件的框架和指导方针来规范提示设计,确保行为的一致性和透明度。
This study addresses the challenges of multi-source feature imbalance and scenario heterogeneity in beamforming optimization for 6G Internet of Things (IoT) networks. To tackle these issues, the work proposes a robust, imbalance-aware optimization framework that integrates multi-perspective features—including network, environmental, device-specific, and visual cues—through a hybrid approach combining supervised and unsupervised learning. For the first time, it systematically evaluates the contribution of each feature type to beamforming performance, revealing that network-related features dominate prediction accuracy, while deployment environment and device type are pivotal for scenario clustering. Scenario segmentation is performed using K-means, DBSCAN, and hierarchical clustering, with model efficacy validated through comprehensive metrics and interpretability analysis. Experiments identify bandwidth, IoT sensor type, and mobility as globally critical features, demonstrating that the proposed method significantly enhances prediction robustness and scenario adaptability.
为评估大型语言模型在高风险心理健康对话中的安全性,开发了K-Bench基准测试,通过临床校准和多情景模拟来评价模型表现。
本文提出了一种基于双阶段MLP-QRNN层次结构的分位数引导特征提取框架,以解决工业制造系统中多时间尺度预测性维护问题。
该论文提出TQRNN30d框架,通过结合量化回归神经网络与多流时间融合分类器,解决长期工业预测维护中的设备退化识别问题。
论文针对LLM在社交机器人中使用时出现的不一致人格问题,通过提出包含八个功能组件的框架和指导方针来规范提示设计,确保行为的一致性和透明度。
This study addresses the challenges of multi-source feature imbalance and scenario heterogeneity in beamforming optimization for 6G Internet of Things (IoT) networks. To tackle these issues, the work proposes a robust, imbalance-aware optimization framework that integrates multi-perspective features—including network, environmental, device-specific, and visual cues—through a hybrid approach combining supervised and unsupervised learning. For the first time, it systematically evaluates the contribution of each feature type to beamforming performance, revealing that network-related features dominate prediction accuracy, while deployment environment and device type are pivotal for scenario clustering. Scenario segmentation is performed using K-means, DBSCAN, and hierarchical clustering, with model efficacy validated through comprehensive metrics and interpretability analysis. Experiments identify bandwidth, IoT sensor type, and mobility as globally critical features, demonstrating that the proposed method significantly enhances prediction robustness and scenario adaptability.