When Is a Steerable Concept Representation Real? Measurement Confounds in a Cross-Family Audit of Neuroscience Parallels in LLMs

πŸ“… 2026-08-08
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenge of disentangling whether human-like neurocognitive traits in large language models stem from intrinsic model properties or artifacts of measurement methodologyβ€”a limitation prevalent in prior work due to reliance on single models and sensitive probing techniques. The authors present the first systematic, cross-model, multi-scale audit across 17 models (0.6B–72B parameters) spanning five model families, integrating linear probing, activation manipulation, residual norm interventions, cross-lingual attribution, and diverse decoding metrics to conduct both causal and correlational analyses. Findings reveal that while concept manipulation effects are widespread, they exhibit no consistent scaling trends; geographic and numerical representations show relative robustness; yet neuron response patterns and cross-lingual structures are highly sensitive to measurement choices. The work underscores the critical need for standardized, controlled evaluation protocols and proposes new criteria for calibrated interventions and operational point selection.
πŸ“ Abstract
Large language models (LLMs) are increasingly reported to exhibit human-like neural and cognitive signatures, including concept cells, mental number lines, and cognitive maps. These claims often rely on linear probing and activation steering applied to a single model, yet both methods are highly sensitive to measurement choices. A reported parallel may therefore reflect the model, the measurement procedure, or both. We audit four representative neuroscience-inspired paradigms across 17 models from five families, spanning $0.6$B to $72$B parameters. Our main experiment examines the causal steerability of concept directions. With raw activation units and a fixed layer and coefficient, steerability appears to increase with model scale, resembling an emergent capability. However, this pattern is produced by an uncalibrated pipeline rather than by a claim established in the steering literature. The trend depends jointly on raw units, the readout metric, and the operating point; correcting any one of these removes it. With residual-norm-comparable interventions and held-out operating-point selection, concept steering remains significant at every scale, but shows no significant trend across the Qwen3 series, although the confidence interval does not rule out a moderate positive slope. The remaining results are mixed. A linear geographic world map is consistently decodable in every tested checkpoint up to $72$B. Number magnitude is strongly encoded, but whether individual neurons appear bell-shaped or monotonic depends on the selection criterion. Language-specific structure is localizable, but the direction of the cross-lingual asymmetry reverses under a different attribution method. These results suggest that the main constraint on AI neuroscience is not a lack of phenomena, but a lack of comparable measurements and adequate controls. We release the protocol, stimuli, and code.
Problem

Research questions and friction points this paper is trying to address.

steerable concept representation
measurement confounds
neuroscience parallels
large language models
activation steering
Innovation

Methods, ideas, or system contributions that make the work stand out.

steerability
measurement confounds
cross-model audit
residual-norm-comparable interventions
concept representation