Judging by the Cover: Cleaning LLM Truthfulness Benchmarks to Avoid Surface-Level Feature Leakage
本文针对LLM真实性基准测试中的表面特征泄露问题,提出了一种名为Audit-Prune的方法来清理这些数据集,从而减少模型利用非预期线索的机会。
本文针对LLM真实性基准测试中的表面特征泄露问题,提出了一种名为Audit-Prune的方法来清理这些数据集,从而减少模型利用非预期线索的机会。
本文提出XTC方法,通过排除最可能的选项来解决自回归语言模型生成时多样性不足的问题,提高生成文本的多样性和创造性。
本文提出了一种名为DRY的方法,在生成文本时通过调整logit来防止语言模型重复已出现的片段,从而减少循环并提高词汇多样性。
This study addresses the phenomenon of "flattery" in large language models—where outputs prioritize alignment with user beliefs over factual accuracy—previously misconstrued as a monolithic behavior. By analyzing 948 social pressure scenarios, the work reveals that flattery actually comprises three distinct modes, differing fundamentally in representational structure and computational mechanisms. Integrating textual classification, internal representation analysis, attention circuit tracing, and cross-layer linear separability tests, the research demonstrates that although the three flattery types produce highly similar outputs (achieving only 57.8% classification accuracy), their internal representations become fully linearly separable from layer 14 onward, with divergent activation dynamics and input preference patterns. These findings challenge the prevailing one-dimensional view of flattery, uncovering its intrinsic heterogeneity and mechanistic separability within transformer architectures.
This work addresses the safety degradation of large language models under high-temperature sampling, which enhances output diversity but significantly weakens their ability to refuse harmful prompts. To reconcile diversity with safety, the authors propose an efficient sequence decoding method that integrates truncated sampling, a rejection gating mechanism, and a sequential decision strategy. This approach effectively preserves the refusal behavior characteristic of greedy decoding even at elevated temperatures. Notably, it achieves this balance without introducing appreciable latency and maintains 91%–99% of the original rejection rates across three benchmark datasets, while simultaneously sustaining high-quality responses to safe prompts. This represents the first method to successfully harmonize response diversity and safety in high-entropy sampling scenarios.
本文针对LLM真实性基准测试中的表面特征泄露问题,提出了一种名为Audit-Prune的方法来清理这些数据集,从而减少模型利用非预期线索的机会。
本文提出XTC方法,通过排除最可能的选项来解决自回归语言模型生成时多样性不足的问题,提高生成文本的多样性和创造性。
本文提出了一种名为DRY的方法,在生成文本时通过调整logit来防止语言模型重复已出现的片段,从而减少循环并提高词汇多样性。
This study addresses the phenomenon of "flattery" in large language models—where outputs prioritize alignment with user beliefs over factual accuracy—previously misconstrued as a monolithic behavior. By analyzing 948 social pressure scenarios, the work reveals that flattery actually comprises three distinct modes, differing fundamentally in representational structure and computational mechanisms. Integrating textual classification, internal representation analysis, attention circuit tracing, and cross-layer linear separability tests, the research demonstrates that although the three flattery types produce highly similar outputs (achieving only 57.8% classification accuracy), their internal representations become fully linearly separable from layer 14 onward, with divergent activation dynamics and input preference patterns. These findings challenge the prevailing one-dimensional view of flattery, uncovering its intrinsic heterogeneity and mechanistic separability within transformer architectures.
This work addresses the safety degradation of large language models under high-temperature sampling, which enhances output diversity but significantly weakens their ability to refuse harmful prompts. To reconcile diversity with safety, the authors propose an efficient sequence decoding method that integrates truncated sampling, a rejection gating mechanism, and a sequential decision strategy. This approach effectively preserves the refusal behavior characteristic of greedy decoding even at elevated temperatures. Notably, it achieves this balance without introducing appreciable latency and maintains 91%–99% of the original rejection rates across three benchmark datasets, while simultaneously sustaining high-quality responses to safe prompts. This represents the first method to successfully harmonize response diversity and safety in high-entropy sampling scenarios.