Language-Representability: Possibilities and Limitations
研究通过扩展二进制语言框架,探讨了小语言如何描述已知图类,并指出该方法在表征稀疏图(如平面图)时的局限性。
研究通过扩展二进制语言框架,探讨了小语言如何描述已知图类,并指出该方法在表征稀疏图(如平面图)时的局限性。
This work addresses the trade-off between granularity and context in factuality evaluation of large language models: atomic facts often lack contextual nuance, while broad statements resist fine-grained assessment. To reconcile this, the authors propose TriQua, a framework that adaptively models factual claims according to their complexity—representing simple assertions as standard triples and encoding complex ones as hyper-relational facts enriched with contextual qualifiers. This approach preserves atomicity while retaining essential context, enabling interpretable, fine-grained error localization. The framework further introduces TriQuaScore, a metric for quantifying factuality at the level of structured factual units. Experiments demonstrate that TriQua achieves robust performance in decomposition quality and alignment with human annotations, outperforming existing methods on evidence-based fact verification tasks.
This work addresses the memory bottleneck of key-value (KV) cache in large language models during long-context inference by introducing FlashJoLT, the first framework to model the KV cache as a third-order tensor. FlashJoLT integrates partial Tucker decomposition to compress token and feature dimensions, Johnson–Lindenstrauss random projections combined with low-bit quantization to preserve residual information, and a Lagrangian dual formulation to jointly optimize compression ratio and reconstruction accuracy. Randomized SVD further accelerates tensor decomposition. Evaluated on Mistral-7B and LLaMA-2-13B, FlashJoLT achieves 2–3× KV cache compression with relative Frobenius errors as low as 0.006–0.009, while maintaining perplexity, GSM8K accuracy, and RULER retrieval performance on par with baselines, and accelerating decomposition by 5–13×.
This study investigates the computational power of two-head finite automata in two-dimensional picture language recognition and their relationship with context-free matrix grammars (CFMGs) and returning pushdown automata (RPDAs). It introduces two novel models of two-head returning finite automata operating on rectangular pictures: the 2-HRFA, which allows backward moves, and the B2-HRFA, which enforces synchronous head movement—a constraint newly introduced in this work. Through formal language-theoretic analysis and closure properties, the paper establishes that the class of languages recognized by 2-HRFA is a proper subset of those recognized by RPDAs and incomparable with CFMGs. Furthermore, it demonstrates that B2-HRFA strictly lies between RFA and 2-HRFA in recognition power, thereby establishing a strict hierarchy among these three models.
This study addresses whether the outputs of small language models (SLMs) in psychometric tasks stem from genuine semantic reasoning or are primarily driven by artifacts of prompt formulation. The authors propose the first diagnostic framework capable of disentangling the influence of such prompt artifacts, systematically manipulating role framing, instructions, item content, and option labels while employing controlled experiments and variance decomposition techniques to quantify the relative contributions of semantic signals versus prompt-induced artifacts. Findings reveal that prompt artifacts frequently dominate model responses, substantially undermining their psychometric validity. The proposed framework not only effectively identifies these confounding influences but also offers a novel pathway for evaluating and enhancing the semantic comprehension capabilities of large language models.
研究通过扩展二进制语言框架,探讨了小语言如何描述已知图类,并指出该方法在表征稀疏图(如平面图)时的局限性。
This work addresses the trade-off between granularity and context in factuality evaluation of large language models: atomic facts often lack contextual nuance, while broad statements resist fine-grained assessment. To reconcile this, the authors propose TriQua, a framework that adaptively models factual claims according to their complexity—representing simple assertions as standard triples and encoding complex ones as hyper-relational facts enriched with contextual qualifiers. This approach preserves atomicity while retaining essential context, enabling interpretable, fine-grained error localization. The framework further introduces TriQuaScore, a metric for quantifying factuality at the level of structured factual units. Experiments demonstrate that TriQua achieves robust performance in decomposition quality and alignment with human annotations, outperforming existing methods on evidence-based fact verification tasks.
This work addresses the memory bottleneck of key-value (KV) cache in large language models during long-context inference by introducing FlashJoLT, the first framework to model the KV cache as a third-order tensor. FlashJoLT integrates partial Tucker decomposition to compress token and feature dimensions, Johnson–Lindenstrauss random projections combined with low-bit quantization to preserve residual information, and a Lagrangian dual formulation to jointly optimize compression ratio and reconstruction accuracy. Randomized SVD further accelerates tensor decomposition. Evaluated on Mistral-7B and LLaMA-2-13B, FlashJoLT achieves 2–3× KV cache compression with relative Frobenius errors as low as 0.006–0.009, while maintaining perplexity, GSM8K accuracy, and RULER retrieval performance on par with baselines, and accelerating decomposition by 5–13×.
This study investigates the computational power of two-head finite automata in two-dimensional picture language recognition and their relationship with context-free matrix grammars (CFMGs) and returning pushdown automata (RPDAs). It introduces two novel models of two-head returning finite automata operating on rectangular pictures: the 2-HRFA, which allows backward moves, and the B2-HRFA, which enforces synchronous head movement—a constraint newly introduced in this work. Through formal language-theoretic analysis and closure properties, the paper establishes that the class of languages recognized by 2-HRFA is a proper subset of those recognized by RPDAs and incomparable with CFMGs. Furthermore, it demonstrates that B2-HRFA strictly lies between RFA and 2-HRFA in recognition power, thereby establishing a strict hierarchy among these three models.
This study addresses whether the outputs of small language models (SLMs) in psychometric tasks stem from genuine semantic reasoning or are primarily driven by artifacts of prompt formulation. The authors propose the first diagnostic framework capable of disentangling the influence of such prompt artifacts, systematically manipulating role framing, instructions, item content, and option labels while employing controlled experiments and variance decomposition techniques to quantify the relative contributions of semantic signals versus prompt-induced artifacts. Findings reveal that prompt artifacts frequently dominate model responses, substantially undermining their psychometric validity. The proposed framework not only effectively identifies these confounding influences but also offers a novel pathway for evaluating and enhancing the semantic comprehension capabilities of large language models.