Active Adaptation, Not Static Defense: Temporal Dynamics of Preventative Steering in Adversarial Fine-Tuning
研究解决大型语言模型对抗恶意微调的脆弱性问题,通过分析预防性引导的时间优化动态,并提出渐进强度调度方法增强防护效果。
研究解决大型语言模型对抗恶意微调的脆弱性问题,通过分析预防性引导的时间优化动态,并提出渐进强度调度方法增强防护效果。
为解决链式思维推理中的计算和上下文成本问题,A*-Thought-V2通过将思维过程建模为隐藏状态轨迹并采用显隐交织的架构,有效压缩偏离主要解题方向的步骤。
Agent memory allows LLM agents to use earlier interactions when answering new queries. Existing methods often compress interaction histories into summaries or other LLM-generated representations. Repeated generation adds cost and can discard answer-bearing details before the system knows what a future query will require. We propose EdgeMem, an agent-memory method built around a simple principle: preserve original interaction turns and organize them through complementary content, temporal, and episodic cues. EdgeMem realizes this principle with a multi-anchor hypergraph constructed by lightweight local processing. Retrieval directly returns source evidence and reserves LLM use for final answer generation, combining structured access to multi-session histories with faithful retention of the original conversation. Experiments on LoCoMo and LongMemEval-S show strong retrieval and memory-grounded question answering; on LoCoMo, EdgeMem achieves the highest strict-judge score among seven reproduced systems under a shared prompt (61.01 versus 58.70), while construction and retrieval require no generative-LLM calls. Overall, EdgeMem shows that preserving and organizing source evidence provides an effective and efficient foundation for agent memory without generative memory management.
为解决移动代理评估中的上下文过载及安全性忽视问题,提出CRATE框架,通过步骤级后果推理和聚合实现任务完成度与操作安全性的自动化评价。
为解决大型音频-语言模型在文化多样性音乐理解上的不足,研究通过构建包含5042个问答对的基准测试集UniVerseBench及自动化生成的训练数据集UniVerseSet,采用不平衡学习策略改进模型性能。
研究解决大型语言模型对抗恶意微调的脆弱性问题,通过分析预防性引导的时间优化动态,并提出渐进强度调度方法增强防护效果。
为解决链式思维推理中的计算和上下文成本问题,A*-Thought-V2通过将思维过程建模为隐藏状态轨迹并采用显隐交织的架构,有效压缩偏离主要解题方向的步骤。
Agent memory allows LLM agents to use earlier interactions when answering new queries. Existing methods often compress interaction histories into summaries or other LLM-generated representations. Repeated generation adds cost and can discard answer-bearing details before the system knows what a future query will require. We propose EdgeMem, an agent-memory method built around a simple principle: preserve original interaction turns and organize them through complementary content, temporal, and episodic cues. EdgeMem realizes this principle with a multi-anchor hypergraph constructed by lightweight local processing. Retrieval directly returns source evidence and reserves LLM use for final answer generation, combining structured access to multi-session histories with faithful retention of the original conversation. Experiments on LoCoMo and LongMemEval-S show strong retrieval and memory-grounded question answering; on LoCoMo, EdgeMem achieves the highest strict-judge score among seven reproduced systems under a shared prompt (61.01 versus 58.70), while construction and retrieval require no generative-LLM calls. Overall, EdgeMem shows that preserving and organizing source evidence provides an effective and efficient foundation for agent memory without generative memory management.
为解决移动代理评估中的上下文过载及安全性忽视问题,提出CRATE框架,通过步骤级后果推理和聚合实现任务完成度与操作安全性的自动化评价。
为解决大型音频-语言模型在文化多样性音乐理解上的不足,研究通过构建包含5042个问答对的基准测试集UniVerseBench及自动化生成的训练数据集UniVerseSet,采用不平衡学习策略改进模型性能。