Institution profile

Dankook University

Academic institutionasia · kr
Official website
Research library11linked papers
Opportunities0open roles
Selected work

Representative Papers

Auditing Question-Order Effects in Large Language Models with the QQ Equality: Mechanism Characterization and a Saturation Caveat

Jul 19, 2026

This study investigates whether large language models (LLMs) adhere to the parameter-free QQ equality—a principle derived from human cognition—when answering sequential questions. To this end, it introduces, for the first time, the quantum question (QQ) equality framework from quantum questionnaire theory into LLM auditing, proposing novel metrics such as saturation diagnostics and sequential sensitivity scores. The authors develop a comprehensive auditing pipeline incorporating worst-case robustness envelopes, sampling-based consistency tests, fully balanced label designs, and saturation analyses. Empirical evaluation on open-source instruction-tuned models reveals that, despite passing all standard validity thresholds, most item pairs elicit near-deterministic (saturated) responses, thereby precluding verification of residual contextuality. This finding suggests that current next-token prediction–based probability distributions are ill-suited as a foundation for modeling survey-style responses.

0 citationsRead paper
Recent publications

Latest Papers

Auditing Question-Order Effects in Large Language Models with the QQ Equality: Mechanism Characterization and a Saturation Caveat

Jul 19, 2026

This study investigates whether large language models (LLMs) adhere to the parameter-free QQ equality—a principle derived from human cognition—when answering sequential questions. To this end, it introduces, for the first time, the quantum question (QQ) equality framework from quantum questionnaire theory into LLM auditing, proposing novel metrics such as saturation diagnostics and sequential sensitivity scores. The authors develop a comprehensive auditing pipeline incorporating worst-case robustness envelopes, sampling-based consistency tests, fully balanced label designs, and saturation analyses. Empirical evaluation on open-source instruction-tuned models reveals that, despite passing all standard validity thresholds, most item pairs elicit near-deterministic (saturated) responses, thereby precluding verification of residual contextuality. This finding suggests that current next-token prediction–based probability distributions are ill-suited as a foundation for modeling survey-style responses.

0 citationsRead paper