A Probe Direction Is a Property of Its Prompt

📅 2026-08-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses a critical oversight in existing probing methods for assessing whether language models are aware of being evaluated: the decisive influence of prompt selection on measurement outcomes, which undermines cross-model comparability. Treating prompts as a core component of measurement design, the authors fix task content while systematically varying prompts and employ controlled experiments, probe direction analysis, variance decomposition, and surface-form ablation tests to quantify the contributions of prompts, models, and their interaction to observed scores. Findings reveal that models account for only a small fraction of score variance; prompt choice can reverse apparent scaling trends; and surface-level prompt features alone suffice to reproduce most published results. The work demonstrates that single-prompt probing yields unreliable comparisons and specifies the minimum number of prompts required for robust evaluation.
📝 Abstract
A model that behaves differently when it senses it is being tested would undermine the evaluations we rely on, so recent work has sought to read that sense directly from a model's activations. The standard instrument contrasts activations on prompts that announce an evaluation against prompts that do not, and reports how well the resulting direction separates held-out cases. That number is then compared across models and correlated with scale. We observe that the instrument has a free parameter its readings do not disclose: "a prompt that announces an evaluation" is not a prompt but a choice among many, and nothing in the method fixes which. Holding the task text fixed and varying only that choice, we find that the reported score, and even the direction in which it trends with model size, follows the prompt rather than the model; two published studies that disagree about the sign of that trend are both reproducible from a single design, by choice of prompt alone. Treating the prompt as a facet of a measurement design rather than an implementation detail, we find the model under study accounts for a small share of the variance in the number reported about it, and most of the rest lies in how each model responds to each prompt: collecting more evaluation items cannot repair the measurement, while varying prompts can. A further check finds that the split these probes are scored on is largely separable from surface form alone, so a direction carrying no information about evaluation at all still reproduces a substantial fraction of each published score. We conclude that a single-prompt design cannot support comparison between models, and we give the number of prompts a defensible comparison requires.
Problem

Research questions and friction points this paper is trying to address.

probe direction
evaluation sensing
prompt dependence
model comparison
measurement variance
Innovation

Methods, ideas, or system contributions that make the work stand out.

probe direction
prompt sensitivity
model evaluation
measurement design
activation probing
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
V
Valentin Noël
Devoteam