KC-Bench: A Dynamic Interactive Benchmark for Evaluating Knowledge Conflicts in LLM Agents
研究提出KC-Bench,通过模拟环境和多源冲突任务评估大语言模型在知识冲突、输入不一致及时间冲突中的处理能力。
研究提出KC-Bench,通过模拟环境和多源冲突任务评估大语言模型在知识冲突、输入不一致及时间冲突中的处理能力。
研究通过分析状态-动作表示、预训练多样性、辅助共训练目标和目标实体暴露四个因素,解决了零样本跨实体VLA传输问题,以提高桌面操作的泛化能力。
This work presents the first unified formalization in Lean 4 that seamlessly connects Scarf’s theorem, Brouwer’s fixed-point theorem, and the existence of mixed Nash equilibria in finite games within a single combinatorial proof framework. By leveraging an indexed-order formulation of Scarf’s theorem, room–door structures, parity arguments, and explicit embedding–projection constructions—combined with compactness and continuity reasoning—the study establishes a rigorous derivation from triangulated simplices to product spaces of simplices. The project not only delivers fully formalized combinatorial proofs of these three foundational results but also introduces BrouwerBench, a benchmark comprising 80 tasks designed to evaluate formal proof systems’ capacity to understand and reason about deep mathematical structures.
This work addresses the challenges of time series forecasting under test-time distribution shifts, where sparse or noisy observation prefixes lead to weak identifiability, error accumulation, and unstable long-term correction. To this end, it formulates test-time adaptation for the first time as a Dirichlet boundary value problem on a temporal manifold. By treating known prefix errors as boundary conditions, the approach combines a local solver for error propagation with a global solver that retrieves cross-window error memory, and introduces Spatio-Temporal Manifold Fusion (SMF) to generate a smooth, bounded correction field. Evaluated across six benchmark datasets and four frozen backbones, the method achieves an average 26.82% relative reduction in MSE over standard baselines and improves upon the strongest baseline by 12.77%, demonstrating remarkable robustness under sparse and corrupted prefix conditions.
This study addresses the challenge of modeling how the quantum yield of fluorescent proteins is governed by their chromophore and its local three-dimensional microenvironment, a task hindered by the difficulty of capturing region-specific physical signals with existing methods. To overcome this, the authors propose a chromophore-centered, typed 3D residue graph representation that incorporates spatial partitioning and a channel–signal–region propagation mechanism, enabling interpretable modeling of edge-specific physical signals and directly revealing wavelength-dependent interaction mechanisms. Coupled with non-identity feature filtering and ExtraTrees regression, the approach achieves a cross-validated R value of 0.772 on a benchmark set of 531 proteins, significantly outperforming current state-of-the-art methods—particularly excelling in tasks involving distantly related homologs (<50% sequence identity) and high-brightness screening.
研究提出KC-Bench,通过模拟环境和多源冲突任务评估大语言模型在知识冲突、输入不一致及时间冲突中的处理能力。
研究通过分析状态-动作表示、预训练多样性、辅助共训练目标和目标实体暴露四个因素,解决了零样本跨实体VLA传输问题,以提高桌面操作的泛化能力。
This work presents the first unified formalization in Lean 4 that seamlessly connects Scarf’s theorem, Brouwer’s fixed-point theorem, and the existence of mixed Nash equilibria in finite games within a single combinatorial proof framework. By leveraging an indexed-order formulation of Scarf’s theorem, room–door structures, parity arguments, and explicit embedding–projection constructions—combined with compactness and continuity reasoning—the study establishes a rigorous derivation from triangulated simplices to product spaces of simplices. The project not only delivers fully formalized combinatorial proofs of these three foundational results but also introduces BrouwerBench, a benchmark comprising 80 tasks designed to evaluate formal proof systems’ capacity to understand and reason about deep mathematical structures.
This work addresses the challenges of time series forecasting under test-time distribution shifts, where sparse or noisy observation prefixes lead to weak identifiability, error accumulation, and unstable long-term correction. To this end, it formulates test-time adaptation for the first time as a Dirichlet boundary value problem on a temporal manifold. By treating known prefix errors as boundary conditions, the approach combines a local solver for error propagation with a global solver that retrieves cross-window error memory, and introduces Spatio-Temporal Manifold Fusion (SMF) to generate a smooth, bounded correction field. Evaluated across six benchmark datasets and four frozen backbones, the method achieves an average 26.82% relative reduction in MSE over standard baselines and improves upon the strongest baseline by 12.77%, demonstrating remarkable robustness under sparse and corrupted prefix conditions.
This study addresses the challenge of modeling how the quantum yield of fluorescent proteins is governed by their chromophore and its local three-dimensional microenvironment, a task hindered by the difficulty of capturing region-specific physical signals with existing methods. To overcome this, the authors propose a chromophore-centered, typed 3D residue graph representation that incorporates spatial partitioning and a channel–signal–region propagation mechanism, enabling interpretable modeling of edge-specific physical signals and directly revealing wavelength-dependent interaction mechanisms. Coupled with non-identity feature filtering and ExtraTrees regression, the approach achieves a cross-validated R value of 0.772 on a benchmark set of 531 proteins, significantly outperforming current state-of-the-art methods—particularly excelling in tasks involving distantly related homologs (<50% sequence identity) and high-brightness screening.