MAxBench: A Multinomial Concept Recovery Benchmark

📅 2026-09-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文引入MAxBench框架,通过对比多种方法来解决多分类概念表示的恢复问题,揭示了不同几何类型在概念控制中的有效性。
📝 Abstract
Fine-grained control of language model behaviors (e.g., steering) is among the more actionable outcomes of interpretability research. For binary concepts such as refusal, a single direction in activation space often suffices for steering. However, many concepts are not binary: Animals and Countries contain many subcategories, each with multiple instances. For these concepts, the search space over possible representation geometries is far larger than for binary concepts; it is thus not clear what geometries are most appropriate, nor what methods are most effective at recovering them. In this work, we introduce MAxBench, a geometry-agnostic evaluation framework for multinomial concept representations based on sampling from the recovered concept representation. We use MAxBench to compare 10 localization methods (covering 5 geometry types) across 6 concepts and 4 models. Using this framework, we find that (i) affine subspaces steer more reliably and have greater recall than rank-one or linear subspaces; (ii) much of this advantage is due to better non-zero offsets rather than the choice of bases; (iii) manifold steering is competitive with the best methods when applicable; and (iv) no method consistently outperforms prompting, in alignment with prior findings on binary concepts. These findings underscore the importance of expanding the scope of interpretability research and meta-evaluation to concepts with more varied structure.
Problem

Research questions and friction points this paper is trying to address.

multinomial concept
representation geometry
concept recovery
activation space
Innovation

Methods, ideas, or system contributions that make the work stand out.

MAxBench
multinomial concept representations
geometry-agnostic evaluation
affine subspaces
manifold steering
🔎 Similar Papers
2024-06-28Conference on Empirical Methods in Natural Language ProcessingCitations: 3