Measuring Activation Control in Large Language Models

📅 2026-08-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过引入激活可控性基准来量化模型通过自然语言指令调节其残差流的能力,以解决大型语言模型在潜空间监控中的潜在欺骗问题。
📝 Abstract
Safe deployment of increasingly capable models will likely come to rely on latent-space monitoring as a complement to behavioral evaluations, especially when evaluation-aware models exhibit scheming or deception. However, if models can also control their own activations, deception could extend into the latent space itself. With this in mind, we introduce the Activation Controllability Benchmark to quantify the extent to which models can modulate their residual stream via natural-language instruction. Across model families and capability levels, we find that most LLMs can control the direction and magnitude of their residual stream activations with some degree of temporal resolution, though performance varies considerably across models. In simple tasks, this level of control can evade activation-based monitoring methods (including linear probes, natural language autoencoders, activation oracles, and the Jacobian lens), albeit imperfectly. These results suggest that control over the activation space itself could become a confound for monitoring as introspective capabilities increase; therefore, we recommend that frontier labs and evaluators track activation controllability in future models.
Problem

Research questions and friction points this paper is trying to address.

Activation Control
Large Language Models
Latent-space Monitoring
Deception
Residual Stream
Innovation

Methods, ideas, or system contributions that make the work stand out.

Activation Controllability Benchmark
residual stream activations
natural-language instruction
latent-space monitoring
🔎 Similar Papers
No similar papers found.