What Attention Recalls and Recurrence Controls in Hybrid Language Models
研究通过两种干预方法探讨了混合语言模型中注意力和循环状态的角色,发现注意力负责精确检索,而循环状态控制输出语言和个性。
研究通过两种干预方法探讨了混合语言模型中注意力和循环状态的角色,发现注意力负责精确检索,而循环状态控制输出语言和个性。
为解决LLM在高风险场景中的事实性问题,提出Enoki框架,通过多级幻觉检测方法,结合文本锚定关系事实的提取与验证,实现高效的事实核验和定位。
研究解决了工具增强型大语言模型在工具返回与参数记忆冲突时的仲裁问题,通过引入MemToC基准进行评估,并采用SFT和DPO方法改进了正确答案保留率。
This study addresses the limitation of existing large language model unlearning methods that overlook fact popularity, rendering high-frequency knowledge difficult to remove. We propose AdaPop, a novel approach that models popularity as a learnable parameter by integrating token confidence with external proxy evaluations. Through a bi-ascent controller, AdaPop dynamically adjusts penalty intensity to achieve an adaptive balance between forgetting and retention. Experiments across three model families and two benchmarks demonstrate that AdaPop reduces content leakage under paraphrased queries by approximately fivefold and under adversarial reconstruction by 1.6 times. Furthermore, internal representation analysis confirms superior forgetting separation, effectively resolving the challenge of unlearning high-frequency facts while preserving general model utility.
This study addresses the ill-posed inverse design problem of V-beam thermal actuators, where multiple geometric and material configurations can yield the same target displacement at a given temperature. To overcome this non-uniqueness, the authors propose a data-driven two-stage optimization framework. First, a neural network forward model is trained to predict thermo-mechanical responses from geometric and material parameters. This model is then frozen and embedded within a gradient-based inverse optimization loop that simultaneously minimizes structural volume and mechanical stress. The approach effectively circumvents the failure of direct regression in ill-conditioned inverse problems and enables efficient multi-objective geometric optimization. Evaluated on a dataset of 3,000 samples, the forward model achieves a mean absolute percentage error (MAPE) of 4.76% in displacement prediction, with over 70% of samples exhibiting errors below 5%.
研究通过两种干预方法探讨了混合语言模型中注意力和循环状态的角色,发现注意力负责精确检索,而循环状态控制输出语言和个性。
为解决LLM在高风险场景中的事实性问题,提出Enoki框架,通过多级幻觉检测方法,结合文本锚定关系事实的提取与验证,实现高效的事实核验和定位。
研究解决了工具增强型大语言模型在工具返回与参数记忆冲突时的仲裁问题,通过引入MemToC基准进行评估,并采用SFT和DPO方法改进了正确答案保留率。
This study addresses the limitation of existing large language model unlearning methods that overlook fact popularity, rendering high-frequency knowledge difficult to remove. We propose AdaPop, a novel approach that models popularity as a learnable parameter by integrating token confidence with external proxy evaluations. Through a bi-ascent controller, AdaPop dynamically adjusts penalty intensity to achieve an adaptive balance between forgetting and retention. Experiments across three model families and two benchmarks demonstrate that AdaPop reduces content leakage under paraphrased queries by approximately fivefold and under adversarial reconstruction by 1.6 times. Furthermore, internal representation analysis confirms superior forgetting separation, effectively resolving the challenge of unlearning high-frequency facts while preserving general model utility.
This study addresses the ill-posed inverse design problem of V-beam thermal actuators, where multiple geometric and material configurations can yield the same target displacement at a given temperature. To overcome this non-uniqueness, the authors propose a data-driven two-stage optimization framework. First, a neural network forward model is trained to predict thermo-mechanical responses from geometric and material parameters. This model is then frozen and embedded within a gradient-based inverse optimization loop that simultaneously minimizes structural volume and mechanical stress. The approach effectively circumvents the failure of direct regression in ill-conditioned inverse problems and enables efficient multi-objective geometric optimization. Evaluated on a dataset of 3,000 samples, the forward model achieves a mean absolute percentage error (MAPE) of 4.76% in displacement prediction, with over 70% of samples exhibiting errors below 5%.