Two Confounds in Cross-Model Value Comparison: Response Determinism and the Access Harness

📅 2026-07-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the confounding of genuine value differences with variations in response determinism in cross-model value comparisons, which is further exacerbated by interference from evaluation harnesses, leading to mischaracterizations of models' individualized values. To disentangle these effects, the authors propose a determinism-corrected decomposition framework that leverages rule-free value dilemmas, repeated forced-choice experiments, and a determinism index to isolate true value divergence from determinism-related artifacts. The approach also systematically evaluates the impact of deployment interfaces—such as APIs and client-side implementations—on value expression. Experiments across nine mainstream language models reveal substantial inter-model differences in determinism (ranging from 0.66 to 0.95) and show that correcting for determinism markedly reduces apparent individualization. Notably, different interfaces induce value profile shifts of up to 0.31 and can even reverse specific moral judgments, thereby uncovering—for the first time—the formative role of the deployment layer in shaping model-expressed values.
📝 Abstract
Cross-model comparisons read divergence in value dispositions as evidence that language models hold individuated values. Under single-draw measurement this conflates two quantities: a difference in central tendency (a genuine value difference) and a difference in response determinism (how sharply a model commits to a forced choice). We introduce a separation protocol -- no-rule value dilemmas with counterbalanced, repeated forced-choice measurement and a determinism index -- and a determinism-corrected decomposition that splits an apparent cross-model distance into a direction-flip component (genuine disagreement) and a same-side-more-extreme component we label determinism. Across nine models, determinism varies substantially (0.66-0.95 among engaging models); whether it is a per-model trait or tracks provider and scale is a question our method makes measurable but our sample leaves open. Correcting for determinism shrinks apparent individuation, while a few cross-family disagreements survive a strict test. We then isolate a second confound: the access harness serving each model. Re-collecting the same models through raw provider APIs, we find the deployment client shifts a model's value profile substantially and client-specifically: one subscription CLI moves a profile by 0.31, flips four of eighteen items, and inflates the flagship's apparent softness (0.34 via CLI vs 0.66 via raw API), whereas another provider's client is clean, confounding provider family with access client. The harness is a value-shaping layer: a base model that refuses one-in-ten forced choices is made compliant by an agent system prompt, established causally in a white-box control. An audit ranking models by single-draw value distance thus ranks a determinism-inflated quantity, confounded further by the client used. We contribute the decomposition and identify the deployment harness as a distinct value confound.
Problem

Research questions and friction points this paper is trying to address.

value comparison
response determinism
access harness
language models
confounding factors
Innovation

Methods, ideas, or system contributions that make the work stand out.

response determinism
access harness
value decomposition
forced-choice measurement
model deployment confound