Auditing Question-Order Effects in Large Language Models with the QQ Equality: Mechanism Characterization and a Saturation Caveat

📅 2026-07-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether large language models (LLMs) adhere to the parameter-free QQ equality—a principle derived from human cognition—when answering sequential questions. To this end, it introduces, for the first time, the quantum question (QQ) equality framework from quantum questionnaire theory into LLM auditing, proposing novel metrics such as saturation diagnostics and sequential sensitivity scores. The authors develop a comprehensive auditing pipeline incorporating worst-case robustness envelopes, sampling-based consistency tests, fully balanced label designs, and saturation analyses. Empirical evaluation on open-source instruction-tuned models reveals that, despite passing all standard validity thresholds, most item pairs elicit near-deterministic (saturated) responses, thereby precluding verification of residual contextuality. This finding suggests that current next-token prediction–based probability distributions are ill-suited as a foundation for modeling survey-style responses.
📝 Abstract
Human survey respondents exhibit question-order effects that satisfy the QQ (quantum question) equality, an a priori, parameter-free prediction of the projective quantum question-order model. We develop the QQ equality into an audit criterion for sequential judgments of autoregressive large language models (LLMs). Theoretically, we characterize which mechanism classes satisfy it robustly: marginal-independent kernels satisfy QQ iff all four mismatch transition rates coincide (a class containing the 2D rank-1 projective model with a fixed measurement pair under state variation); a polarity- and position-dependent repetition family is characterized by an exact cross-symmetry condition with closed-form violations; QQ-satisfying behaviors are closed under order-matched mixing; and the rank-2 Contextuality-by-Default criterion translates into audit coordinates as $|\qQQ|\le\OSS$, where $\OSS$ (the order-sensitivity score) totals the order sensitivity of the two marginals. Methodologically, we develop a pre-specified, audit-logged pipeline applicable to any model exposing next-token log-probabilities; it combines worst-case robustness envelopes, sampling-consistency spot checks, full label counterbalancing, and a saturation diagnostic. Empirically, in a first-signal pilot on an open-weight instruction-tuned model under two framings, all pre-specified health gates passed, yet 17/18 and 7/8 item pairs, respectively, were saturated (near-deterministic), and no item was certified residually contextual. Forced-binary next-token log-probabilities were thus inadequate for distribution-level QQ audits under the tested model and prompting conditions; we recommend pre-specified saturation diagnostics whenever next-token distributions are treated as survey-response distributions.
Problem

Research questions and friction points this paper is trying to address.

question-order effects
QQ equality
large language models
sequential judgments
survey-response distributions
Innovation

Methods, ideas, or system contributions that make the work stand out.

QQ equality
question-order effects
large language models
audit framework
saturation diagnostic