Institution profile

UK AI Security Institute

Academic institutioneurope · gb
Official website
Research library24linked papers
Opportunities0open roles
Selected work

Representative Papers

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii

May 02, 2025

Mechanistic interpretability (MI) lacks unified, actionable causal evaluation criteria, hindering its advancement. This paper addresses this gap by systematically integrating four classical philosophical accounts of explanation—Bayesian, Kuhnian, Deductive-Nomological, and Mechanistic—into the first multidimensional evaluation framework tailored for mechanistic explanations. We introduce the “compact proof” paradigm: a novel explanatory form that jointly satisfies concision, unification, and generality, and formally specify its generation and verification procedures. Empirical evaluation demonstrates that this paradigm significantly improves explanation quality across diverse models and tasks. Beyond resolving MI’s evaluation bottleneck, our work identifies three foundational research directions: formalizing concision, modeling explanatory unification, and deriving domain-general principles. The framework provides theoretical foundations and methodological tools for building trustworthy AI systems that are monitorable, predictable, and controllable.

1 citationsRead paper

Knowing When to Stop: Bayesian Optimal Stopping for LLM Evaluations

Aug 14, 2026

This study addresses the inefficiency of fixed sampling in large language model evaluation, which often leads to resource wastage due to the absence of adaptive stopping mechanisms. We propose OptStop, a framework introducing a novel calibration-free hierarchical Bayesian adaptive stopping strategy that formulates evaluation as a sequential measurement problem. This approach dynamically allocates samples based on uncertainty quantification while incorporating a zero-performance safety fallback. Extensive experiments across nine validation settings demonstrate that OptStop achieves accuracy equivalent to exhaustive evaluation while reducing the required number of trials by 57% to 97%. These results indicate substantial improvements in both evaluation efficiency and computational resource utilization, offering a robust solution for cost-effective LLM benchmarking without compromising assessment reliability.

0 citationsRead paper
Recent publications

Latest Papers

Knowing When to Stop: Bayesian Optimal Stopping for LLM Evaluations

Aug 14, 2026

This study addresses the inefficiency of fixed sampling in large language model evaluation, which often leads to resource wastage due to the absence of adaptive stopping mechanisms. We propose OptStop, a framework introducing a novel calibration-free hierarchical Bayesian adaptive stopping strategy that formulates evaluation as a sequential measurement problem. This approach dynamically allocates samples based on uncertainty quantification while incorporating a zero-performance safety fallback. Extensive experiments across nine validation settings demonstrate that OptStop achieves accuracy equivalent to exhaustive evaluation while reducing the required number of trials by 57% to 97%. These results indicate substantial improvements in both evaluation efficiency and computational resource utilization, offering a robust solution for cost-effective LLM benchmarking without compromising assessment reliability.

0 citationsRead paper

Item Response Theory for AI Safety

Aug 05, 2026

Existing AI safety evaluation benchmarks suffer from redundancy, high inter-correlation, and potential sandbagging—where models deliberately underperform—rendering aggregated scores difficult to interpret and trust. This work presents the first large-scale application of Item Response Theory (IRT) to safety assessment of large language models (LLMs), integrating factor analysis and adaptive testing to distill three interpretable latent traits from multiple benchmarks: refusal strictness, truthfulness, and contextual harm. The proposed IRT-based framework achieves 97–99% fidelity in replicating individual benchmark outcomes using only about ten adaptively selected items, substantially improving evaluation efficiency. Furthermore, it effectively detects sandbagging behaviors, demonstrating IRT’s reliability, scalability, and auditability in LLM safety evaluation.

0 citationsRead paper