CBX-Bench: A Human-Aligned MLLM Council for Benchmarking Concept Bottleneck Model Explanations

📅 2026-08-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of quantitative metrics and ground-truth annotations for evaluating explanations in Concept Bottleneck Models (CBMs). We propose CBX-Bench, the first explanation evaluation benchmark aligned with human preferences, alongside a Multimodal Large Language Model (MLLM) review mechanism. By constructing an MLLM ensemble jury and a comparative evaluation framework, this work achieves scalable, cognitively consistent automated scoring. Experimental results demonstrate that the proposed jury recovers over 70% of strict human preference rankings with an 83% agreement rate. Furthermore, we establish a public leaderboard to bridge the gap in quantitative interpretability assessment for CBMs. Collectively, these contributions provide a reliable benchmark for trustworthy AI research, facilitating standardized evaluation of model explanations where traditional metrics have previously fallen short.
📝 Abstract
Concept Bottleneck Models (CBMs) are designed to make visual classification interpretable by expressing predictions through human-understandable concepts. Although interpretability is the central motivation for CBMs, they are still largely evaluated as predictive models by downstream classification accuracy, supplemented by isolated qualitative examples. This highlights a pressing need for quantitative measures, a challenge complicated by the infeasibility of ground-truth concept annotation at scale and the open nature of concept lists due to a lack of consensus. To fill this gap, we develop a multimodal large language model (MLLM) council that, given an image and its CBM explanation, produces an explanation quality score. To ground and validate the council, we first conduct a human study to establish a ground-truth reference for CBM explanation quality: for an image, annotators compare explanations from two of LF-CBM, VLG-CBM, and CBM-Suite and choose the more useful one, or mark them as equally good or equally bad, yielding 2700 judgments over 900 image-comparison items on CUB-200, ImageNet-100, and Places365. Against this human reference, our five-model council, consisting of open-weight MLLMs, recovers over 70% of strict human preference rankings, rising to 83% on items where human annotators unanimously agree. Building on this validated council, we introduce CBX-Bench, a public benchmark and leaderboard: authors of new CBMs can submit their model's explanations, and CBX-Bench scores them with the council and maintains dataset-level rankings of explanation quality. CBX-Bench thus provides a human-aligned, scalable evaluation of CBM explanations beyond accuracy and isolated qualitative examples. The benchmark is available at https://github.com/meric-karadag/cbx-bench.
Problem

Research questions and friction points this paper is trying to address.

Concept Bottleneck Models
Explainability Evaluation
Quantitative Metrics
Human Alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Concept Bottleneck Models
Multimodal Large Language Model Council
Explanation Quality Evaluation
Human-Aligned Benchmark
CBX-Bench
🔎 Similar Papers
No similar papers found.
Y
Yusuf Meric Karadag
Department of Computer Engineering, Middle East Technical University (METU)
G
Gulay Oklan
Department of Computer Engineering, Middle East Technical University (METU)
S
Seref Baris Cagliyan
Department of Computer Engineering, Middle East Technical University (METU)
U
Umut Ozdemir
Department of Computer Engineering, Middle East Technical University (METU)
Emre Akbas
Emre Akbas
Helmholtz Munich | Middle East Technical University (METU)
computer visiondeep learningmachine learningobject detectionhuman pose estimation