When Decodability Is Not Enough: Logical Validity Representations, Behavioral Dissociation, and Causal Tests in Language Models

📅 2026-09-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究探讨了五个开放权重转换模型在逻辑验证中的内部表示问题,通过使用有效-无效前提-声明对进行测试,发现尽管行为表现接近随机,逻辑有效性仍可从隐藏状态中解码。
📝 Abstract
Large language models can look capable of logical reasoning, but correct or incorrect answers alone tell us little about what the model represents internally. We study logical verification in five open-weight transformer models using matched valid--invalid premise--claim pairs that vary across inference families, semantic domains, templates, and difficulty levels. Despite near-chance behavioral performance, logical validity is often almost perfectly decodable from hidden states and remains strongly decodable under held-out templates, domains, and inference families. Validity also remains highly decodable on behaviorally incorrect examples in the conditions where correctness-conditioned evaluation is well defined. At the same time, exhaustive leave-one-out tests reveal clear limits to this generalization, and interventions along probe-derived validity directions have only weak, nonspecific effects compared with random controls. Our results suggest that representing validity, expressing it in behavior, and using it causally are distinct. Validity related information can be strongly decodable from a model's hidden states without being reliably expressed in its output.
Problem

Research questions and friction points this paper is trying to address.

logical reasoning
internal representation
logical validity
Innovation

Methods, ideas, or system contributions that make the work stand out.

logical validity
hidden states
decodability
transformer models
behavioral expression
🔎 Similar Papers
No similar papers found.