Record Grouping Controls Evidence Weight in Language Models

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过控制记录分组来调整语言模型中的证据权重,提出了一种预生成表示方法以去除组内重复内容并限制每组贡献,实验验证了不同分组对模型决策的影响。
📝 Abstract
Retrieved records are presentation units; a supplied partition determines which records enter a language model as one evidential contribution. We characterize the invariant group-content state that removes within-group copies while retaining complementary canonical content, show that equal group counts can encode different evidence states, and derive a sharp content-aware partition-error bound. Given a supplied partition, our pre-generation representation deduplicates and aggregates content within groups and bounds each group's contribution. Across 104,402 trials and 6 public checkpoints, a central natural-text intervention finds that content-fixed false splits add 10.27-32.66 percentage points and false merges remove 9.13-31.79 points; a matched six-slot control retains the positive direction in all 16 cells. In a new 48-item controlled campaign panel, changing the supplied partition produces measurable, checkpoint-dependent decision shifts across all four models, and the balanced mirror design exposes substantial order interactions. Together, the theory and experiments establish the supplied partition as a controllable pre-generation representation variable and characterize its checkpoint-dependent behavioral effects.
Problem

Research questions and friction points this paper is trying to address.

Record Grouping
Evidence Weight
Language Models
Partition
Content Deduplication
Innovation

Methods, ideas, or system contributions that make the work stand out.

Record Grouping
Evidence Weight
Language Models
Content Deduplication
Partition-Error Bound
Z
Zhongxuan Liu
Faculty of Computing, Harbin Institute of Technology
S
Sicheng Zhou
Faculty of Computing, Harbin Institute of Technology
Hongzhi Wang
Hongzhi Wang
IBM Almaden Research Center
Medical Image Analysis