Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Frontier Language Models
研究审计了22种前沿模型在12个回归基准上直接检索已发表数值的情况,探讨了不同推理层次对检索的影响,并测试了一种中断检索的方法。
研究审计了22种前沿模型在12个回归基准上直接检索已发表数值的情况,探讨了不同推理层次对检索的影响,并测试了一种中断检索的方法。
本文提出了一种基于深度学习图像质量度量的全可微抖动校正方法,用于X射线相衬显微CT,直接从投影数据估计和补偿每个投影的刚性抖动。
本文提出ADDA框架,通过支持自动微分和并行计算解决数据同化中模拟与同化代码不兼容等问题。
This study addresses the limitations of single-path evaluation and neglected anchor relevance in assessing anchoring effects within large language models. We construct a multi-path anchoring benchmark incorporating an explicit relevance dimension and conduct large-scale controlled experiments across fourteen models. Our results reveal the path-dependency of anchoring bias, demonstrating that plausible anchors induce more significant deviations and that high-accuracy models remain vulnerable to such semantically relevant anchors. By transcending traditional single-path evaluation paradigms, this work systematically elucidates the mechanisms underlying model bias across diverse reasoning paths and anchor semantics. Ultimately, these findings establish a novel framework for evaluating the cognitive robustness of large language models, highlighting critical vulnerabilities even in high-performing systems when exposed to contextually plausible misinformation.
In data assimilation, uncertainty quantification remains challenging due to the coupling of dynamical models with process noise and sparse, noisy observations. To address this, we propose a variational inference–based uncertainty-aware data assimilation framework that models stochastic state evolution as a multivariate Gaussian distribution, enabling approximately perfectly calibrated uncertainty estimates and supporting longer assimilation windows. The method is end-to-end differentiable and seamlessly embeddable within machine learning architectures. Evaluated on the chaotic Lorenz-96 system, it achieves high state estimation accuracy while significantly improving uncertainty calibration and out-of-distribution generalization compared to conventional approaches. All code is publicly available to facilitate reproducibility and further research extensions.
研究审计了22种前沿模型在12个回归基准上直接检索已发表数值的情况,探讨了不同推理层次对检索的影响,并测试了一种中断检索的方法。
本文提出了一种基于深度学习图像质量度量的全可微抖动校正方法,用于X射线相衬显微CT,直接从投影数据估计和补偿每个投影的刚性抖动。
本文提出ADDA框架,通过支持自动微分和并行计算解决数据同化中模拟与同化代码不兼容等问题。
This study addresses the limitations of single-path evaluation and neglected anchor relevance in assessing anchoring effects within large language models. We construct a multi-path anchoring benchmark incorporating an explicit relevance dimension and conduct large-scale controlled experiments across fourteen models. Our results reveal the path-dependency of anchoring bias, demonstrating that plausible anchors induce more significant deviations and that high-accuracy models remain vulnerable to such semantically relevant anchors. By transcending traditional single-path evaluation paradigms, this work systematically elucidates the mechanisms underlying model bias across diverse reasoning paths and anchor semantics. Ultimately, these findings establish a novel framework for evaluating the cognitive robustness of large language models, highlighting critical vulnerabilities even in high-performing systems when exposed to contextually plausible misinformation.
In data assimilation, uncertainty quantification remains challenging due to the coupling of dynamical models with process noise and sparse, noisy observations. To address this, we propose a variational inference–based uncertainty-aware data assimilation framework that models stochastic state evolution as a multivariate Gaussian distribution, enabling approximately perfectly calibrated uncertainty estimates and supporting longer assimilation windows. The method is end-to-end differentiable and seamlessly embeddable within machine learning architectures. Evaluated on the chaotic Lorenz-96 system, it achieves high state estimation accuracy while significantly improving uncertainty calibration and out-of-distribution generalization compared to conventional approaches. All code is publicly available to facilitate reproducibility and further research extensions.