On the Impact of Anonymization on the Performance of Large Language Models
研究探讨了数据匿名化对大型语言模型性能的影响,通过对比不同模型在原始与匿名化输入下的表现,揭示了匿名化对模型能力的复杂影响及任务依赖性。
研究探讨了数据匿名化对大型语言模型性能的影响,通过对比不同模型在原始与匿名化输入下的表现,揭示了匿名化对模型能力的复杂影响及任务依赖性。
论文提出了一种基于分布差异的方法来评估量化大型语言模型的保真度损失,通过计算全精度和量化模型之间的统计距离,提供了比仅依赖准确率更细致的分析。
该研究解决了量子LDPC码的解码难题,通过构建概率模型并开发两种新型解码器,一种基于采样提供最优性证明,另一种基于区域实现精确最大似然解码。
This work addresses the limitations of existing audio description generation methods, which struggle with long-form videos and rely on manually annotated timestamps, thereby failing to meet the accessibility needs of visually impaired users at scale. The paper introduces the first streaming framework for audio description generation tailored to long videos, leveraging a sliding window mechanism to enable real-time caption insertion without ground-truth timestamps. The framework supports both fine-tuning (StrAD-FT) and zero-shot prompting with vision-language models (StrAD-Zero). Additionally, the authors construct StrAD, a diverse benchmark of long videos, to standardize full-video-level evaluation. Experiments show that the proposed method achieves a CIDEr score of 36.3 on CMD-AD—outperforming prior work by 10.0—and reaches 51.0 CIDEr on the StrAD benchmark, with a streaming task SODA score of 2.4, substantially exceeding the zero-shot baseline of 1.1.
This work addresses the tendency of existing recurrent Transformers to suffer from overthinking and computational redundancy due to the absence of cross-iteration guidance signals. To mitigate this, the authors propose integrating recurrence with multi-token prediction (MTP), aligning the number of recurrence steps with future token prediction horizons in latent space—such that the t-th recurrence step directly predicts the token t steps ahead—thereby providing dense lookahead supervision. A lightweight gating mechanism is further introduced to preserve useful information across iterations. This approach achieves, for the first time, an explicit alignment between recurrence depth and prediction horizon in latent space, significantly alleviating overthinking and enhancing inference efficiency. Experiments demonstrate up to an 8.1% relative improvement in average accuracy over non-recurrent baselines, with stable training observed even with up to 15 recurrence steps.
研究探讨了数据匿名化对大型语言模型性能的影响,通过对比不同模型在原始与匿名化输入下的表现,揭示了匿名化对模型能力的复杂影响及任务依赖性。
论文提出了一种基于分布差异的方法来评估量化大型语言模型的保真度损失,通过计算全精度和量化模型之间的统计距离,提供了比仅依赖准确率更细致的分析。
该研究解决了量子LDPC码的解码难题,通过构建概率模型并开发两种新型解码器,一种基于采样提供最优性证明,另一种基于区域实现精确最大似然解码。
This work addresses the limitations of existing audio description generation methods, which struggle with long-form videos and rely on manually annotated timestamps, thereby failing to meet the accessibility needs of visually impaired users at scale. The paper introduces the first streaming framework for audio description generation tailored to long videos, leveraging a sliding window mechanism to enable real-time caption insertion without ground-truth timestamps. The framework supports both fine-tuning (StrAD-FT) and zero-shot prompting with vision-language models (StrAD-Zero). Additionally, the authors construct StrAD, a diverse benchmark of long videos, to standardize full-video-level evaluation. Experiments show that the proposed method achieves a CIDEr score of 36.3 on CMD-AD—outperforming prior work by 10.0—and reaches 51.0 CIDEr on the StrAD benchmark, with a streaming task SODA score of 2.4, substantially exceeding the zero-shot baseline of 1.1.
This work addresses the tendency of existing recurrent Transformers to suffer from overthinking and computational redundancy due to the absence of cross-iteration guidance signals. To mitigate this, the authors propose integrating recurrence with multi-token prediction (MTP), aligning the number of recurrence steps with future token prediction horizons in latent space—such that the t-th recurrence step directly predicts the token t steps ahead—thereby providing dense lookahead supervision. A lightweight gating mechanism is further introduced to preserve useful information across iterations. This approach achieves, for the first time, an explicit alignment between recurrence depth and prediction horizon in latent space, significantly alleviating overthinking and enhancing inference efficiency. Experiments demonstrate up to an 8.1% relative improvement in average accuracy over non-recurrent baselines, with stable training observed even with up to 15 recurrence steps.