Institution profile

Kitware, Inc.

Industry researchnorthamerica · us
Official website
Research library16linked papers
Opportunities0open roles
Selected work

Representative Papers

Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results

Jun 12, 2026

This work addresses the challenge of incomparable, non-reusable, and fragmented AI evaluation results stemming from heterogeneous formats, disparate sources, and inconsistent frameworks. To overcome this, the authors propose the first community-governed unified standard that defines JSON Schemas for both metadata and instance-level evaluation results, alongside a source-agnostic architecture and automated conversion tools supporting 31 diverse evaluation formats. Leveraging this standard, they have constructed a large-scale, standardized database encompassing 22,235 models and 2,273 benchmarks, hosted on Hugging Face with support for crowdsourced contributions. This infrastructure significantly enhances the comparability and reusability of evaluation results while fostering efficient cross-community collaboration.

0 citationsRead paper

REBAR: Reference Ethical Benchmark for Autonomy Readiness

May 18, 2026

This study addresses the critical gap in objective, computable metrics for quantifying the ethical and legal compliance of autonomous systems, which has hindered their evaluability and the development of robust accountability mechanisms. To bridge this gap, the authors propose a large language model framework integrating neuro-symbolic methods that maps system behaviors onto an interpretable “Autonomy Readiness Level” (ARL) scale through high-fidelity simulation and automated test generation. This approach enables, for the first time, objective and reproducible benchmark scoring of ethical performance in white-box autonomous systems, effectively closing the divide between abstract ethical principles and verifiable, accountable behaviors.

0 citationsRead paper
Recent publications

Latest Papers

Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results

Jun 12, 2026

This work addresses the challenge of incomparable, non-reusable, and fragmented AI evaluation results stemming from heterogeneous formats, disparate sources, and inconsistent frameworks. To overcome this, the authors propose the first community-governed unified standard that defines JSON Schemas for both metadata and instance-level evaluation results, alongside a source-agnostic architecture and automated conversion tools supporting 31 diverse evaluation formats. Leveraging this standard, they have constructed a large-scale, standardized database encompassing 22,235 models and 2,273 benchmarks, hosted on Hugging Face with support for crowdsourced contributions. This infrastructure significantly enhances the comparability and reusability of evaluation results while fostering efficient cross-community collaboration.

0 citationsRead paper

REBAR: Reference Ethical Benchmark for Autonomy Readiness

May 18, 2026

This study addresses the critical gap in objective, computable metrics for quantifying the ethical and legal compliance of autonomous systems, which has hindered their evaluability and the development of robust accountability mechanisms. To bridge this gap, the authors propose a large language model framework integrating neuro-symbolic methods that maps system behaviors onto an interpretable “Autonomy Readiness Level” (ARL) scale through high-fidelity simulation and automated test generation. This approach enables, for the first time, objective and reproducible benchmark scoring of ethical performance in white-box autonomous systems, effectively closing the divide between abstract ethical principles and verifiable, accountable behaviors.

0 citationsRead paper