UniqueShip: Mitigating Data Leakage in Acoustic Ship Classification Benchmark Datasets
本文通过引入UniqueShip数据集解决水下声学目标识别中数据泄漏问题,采用控制训练与评估集间数据独立性的方法,提高模型评价可靠性。
本文通过引入UniqueShip数据集解决水下声学目标识别中数据泄漏问题,采用控制训练与评估集间数据独立性的方法,提高模型评价可靠性。
研究针对高维生物医学数据的复杂界面问题,通过两组实验评估传统和特定领域的可用性标准,提出三个新启发式方法以改善多学科研究平台的用户体验。
本文提出一种架构,通过分离意图解释、执行和解释,并基于领域本体约束分析链,解决了LLM辅助科学可视化中生成错误脚本的问题。
This work addresses the challenge of incomparable, non-reusable, and fragmented AI evaluation results stemming from heterogeneous formats, disparate sources, and inconsistent frameworks. To overcome this, the authors propose the first community-governed unified standard that defines JSON Schemas for both metadata and instance-level evaluation results, alongside a source-agnostic architecture and automated conversion tools supporting 31 diverse evaluation formats. Leveraging this standard, they have constructed a large-scale, standardized database encompassing 22,235 models and 2,273 benchmarks, hosted on Hugging Face with support for crowdsourced contributions. This infrastructure significantly enhances the comparability and reusability of evaluation results while fostering efficient cross-community collaboration.
This study addresses the critical gap in objective, computable metrics for quantifying the ethical and legal compliance of autonomous systems, which has hindered their evaluability and the development of robust accountability mechanisms. To bridge this gap, the authors propose a large language model framework integrating neuro-symbolic methods that maps system behaviors onto an interpretable “Autonomy Readiness Level” (ARL) scale through high-fidelity simulation and automated test generation. This approach enables, for the first time, objective and reproducible benchmark scoring of ethical performance in white-box autonomous systems, effectively closing the divide between abstract ethical principles and verifiable, accountable behaviors.
本文通过引入UniqueShip数据集解决水下声学目标识别中数据泄漏问题,采用控制训练与评估集间数据独立性的方法,提高模型评价可靠性。
研究针对高维生物医学数据的复杂界面问题,通过两组实验评估传统和特定领域的可用性标准,提出三个新启发式方法以改善多学科研究平台的用户体验。
本文提出一种架构,通过分离意图解释、执行和解释,并基于领域本体约束分析链,解决了LLM辅助科学可视化中生成错误脚本的问题。
This work addresses the challenge of incomparable, non-reusable, and fragmented AI evaluation results stemming from heterogeneous formats, disparate sources, and inconsistent frameworks. To overcome this, the authors propose the first community-governed unified standard that defines JSON Schemas for both metadata and instance-level evaluation results, alongside a source-agnostic architecture and automated conversion tools supporting 31 diverse evaluation formats. Leveraging this standard, they have constructed a large-scale, standardized database encompassing 22,235 models and 2,273 benchmarks, hosted on Hugging Face with support for crowdsourced contributions. This infrastructure significantly enhances the comparability and reusability of evaluation results while fostering efficient cross-community collaboration.
This study addresses the critical gap in objective, computable metrics for quantifying the ethical and legal compliance of autonomous systems, which has hindered their evaluability and the development of robust accountability mechanisms. To bridge this gap, the authors propose a large language model framework integrating neuro-symbolic methods that maps system behaviors onto an interpretable “Autonomy Readiness Level” (ARL) scale through high-fidelity simulation and automated test generation. This approach enables, for the first time, objective and reproducible benchmark scoring of ethical performance in white-box autonomous systems, effectively closing the divide between abstract ethical principles and verifiable, accountable behaviors.