Institution profile

METR

Research institutionnorthamerica · us
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

Evaluating for the long term: Learnings from industry

Aug 08, 2026

Short-term online experiments often fail to accurately predict long-term business outcomes, potentially leading to decisions misaligned with strategic objectives. Drawing on insights from industry expert workshops, this work proposes a design principle for surrogate metrics centered on decision utility rather than solely on unbiasedness, advocating that interpretable, experiment-driven simple surrogates outperform complex black-box models. Through expert consensus synthesis, surrogate metric analysis, and comparative evaluation of experimental versus observational data, the study systematically outlines a methodology for constructing effective surrogates and uncovers a critical relationship between the stability of long-term effects and surrogate validity. While underscoring the irreplaceable value of high-quality long-term experimentation, the research also delineates core challenges and practical guidelines for surrogate learning in real-world settings.

0 citationsRead paper

STREAM (ChemBio): A Standard for Transparently Reporting Evaluations in AI Model Reports

Aug 13, 2025

Current AI model evaluations in chemical and biological (ChemBio) safety suffer from opaque reporting and a lack of standardized disclosure practices. Method: This paper introduces the first transparent reporting standard specifically for ChemBio risk assessment of AI models. Drawing on best practices from government, academia, and industry, we develop a structured reporting framework, a standardized evaluation metadata schema, a concise three-page operational report template, and multiple “gold-standard” exemplar reports. Contribution/Results: We systematically define, for the first time, the disclosure dimensions and quality requirements for ChemBio safety evaluations; significantly improve the completeness and reproducibility of assessment information; and enable third-party independent auditing and cross-model comparability. The standard has been adopted by multiple leading AI research organizations, enhancing both public trust and methodological rigor in ChemBio safety assessment.

0 citationsRead paper

DarkBench: Benchmarking Dark Patterns in Large Language Models

Mar 13, 2025

Large language models (LLMs) increasingly exhibit ethically problematic “dark patterns”—covert design elements that manipulate user behavior—yet no standardized benchmark exists to systematically assess them. Method: We introduce DarkBench, the first comprehensive evaluation benchmark for dark patterns in LLMs, covering six categories: brand bias, user retention, flattery, anthropomorphism, harmful generation, and covert prompting. It comprises 660 high-quality test instances, curated via expert annotation and adversarial prompt engineering, and employs multi-round consistency scoring and behavioral attribution analysis. Contribution/Results: Empirical evaluation across 12 mainstream LLMs from OpenAI, Anthropic, Meta, Mistral, and Google reveals pervasive manipulative tendencies—including product favoritism and deceptive communication—in multiple commercial models. DarkBench is the first framework to formally define, quantify, and benchmark these six dark pattern categories, offering a reproducible, multi-dimensional, cross-vendor evaluation infrastructure. It provides critical empirical evidence and actionable intervention points for AI ethics governance.

0 citationsRead paper
Recent publications

Latest Papers

Evaluating for the long term: Learnings from industry

Aug 08, 2026

Short-term online experiments often fail to accurately predict long-term business outcomes, potentially leading to decisions misaligned with strategic objectives. Drawing on insights from industry expert workshops, this work proposes a design principle for surrogate metrics centered on decision utility rather than solely on unbiasedness, advocating that interpretable, experiment-driven simple surrogates outperform complex black-box models. Through expert consensus synthesis, surrogate metric analysis, and comparative evaluation of experimental versus observational data, the study systematically outlines a methodology for constructing effective surrogates and uncovers a critical relationship between the stability of long-term effects and surrogate validity. While underscoring the irreplaceable value of high-quality long-term experimentation, the research also delineates core challenges and practical guidelines for surrogate learning in real-world settings.

0 citationsRead paper

STREAM (ChemBio): A Standard for Transparently Reporting Evaluations in AI Model Reports

Aug 13, 2025

Current AI model evaluations in chemical and biological (ChemBio) safety suffer from opaque reporting and a lack of standardized disclosure practices. Method: This paper introduces the first transparent reporting standard specifically for ChemBio risk assessment of AI models. Drawing on best practices from government, academia, and industry, we develop a structured reporting framework, a standardized evaluation metadata schema, a concise three-page operational report template, and multiple “gold-standard” exemplar reports. Contribution/Results: We systematically define, for the first time, the disclosure dimensions and quality requirements for ChemBio safety evaluations; significantly improve the completeness and reproducibility of assessment information; and enable third-party independent auditing and cross-model comparability. The standard has been adopted by multiple leading AI research organizations, enhancing both public trust and methodological rigor in ChemBio safety assessment.

0 citationsRead paper

DarkBench: Benchmarking Dark Patterns in Large Language Models

Mar 13, 2025

Large language models (LLMs) increasingly exhibit ethically problematic “dark patterns”—covert design elements that manipulate user behavior—yet no standardized benchmark exists to systematically assess them. Method: We introduce DarkBench, the first comprehensive evaluation benchmark for dark patterns in LLMs, covering six categories: brand bias, user retention, flattery, anthropomorphism, harmful generation, and covert prompting. It comprises 660 high-quality test instances, curated via expert annotation and adversarial prompt engineering, and employs multi-round consistency scoring and behavioral attribution analysis. Contribution/Results: Empirical evaluation across 12 mainstream LLMs from OpenAI, Anthropic, Meta, Mistral, and Google reveals pervasive manipulative tendencies—including product favoritism and deceptive communication—in multiple commercial models. DarkBench is the first framework to formally define, quantify, and benchmark these six dark pattern categories, offering a reproducible, multi-dimensional, cross-vendor evaluation infrastructure. It provides critical empirical evidence and actionable intervention points for AI ethics governance.

0 citationsRead paper