Toward Self-Adaptive Physical AI: Can LLM Agents Manage Long-Horizon Physical Tasks?
研究探索了利用零样本大型语言模型代理执行长期物理任务的可行性,通过设计一个集成了规划、工具调用、观察和验证功能的多代理框架,在农业任务中表现出色,并能更好地适应环境变化。
研究探索了利用零样本大型语言模型代理执行长期物理任务的可行性,通过设计一个集成了规划、工具调用、观察和验证功能的多代理框架,在农业任务中表现出色,并能更好地适应环境变化。
本文提出了一种近似算法,用于计算分段代数曲线之间的连续动态时间规整(CDTW)距离,以解决现有方法对采样率和异常值敏感的问题。
Existing pruning methods struggle to simultaneously preserve task performance and biological plausibility, particularly in functional recurrent neural networks. This work proposes and empirically validates noise-prune, a novel local pruning approach based on synaptic noise fluctuations: it retains critical connections through local sampling and rescales their weights to maintain average synaptic strength. The study demonstrates that both the connection sampling strategy and the weight rescaling mechanism are essential for performance, and revises the theoretically predicted optimal rescaling magnitude. Noise-prune significantly outperforms magnitude-only pruning strategies and matches or even exceeds the performance of non-local methods that rely on second-order information, while remaining computationally efficient and biologically plausible.
Existing reasoning-time strategies exhibit critical limitations: self-correction tends to reinforce initial biases, multi-agent collaboration (MAC) often suffers from insufficient coordination leading to collective errors, and high-accuracy verifiers require extensive human annotations. This paper proposes AdCo, the first framework to introduce a UCB-based adaptive “coopetition” mechanism—dynamically balancing cooperation and competition—into multi-agent LLM reasoning. AdCo leverages only coarse-grained verification signals to guide uncertainty-aware exploration and enhance trajectory diversity. Its methodology integrates multi-agent collaborative reasoning, iterative refinement, reasoning trajectory analysis, and knowledge diversity modeling. Evaluated on multiple mathematical reasoning benchmarks, AdCo achieves a 20% relative performance gain over state-of-the-art baselines while demonstrating robustness across varying sample sizes and configuration settings.
研究探索了利用零样本大型语言模型代理执行长期物理任务的可行性,通过设计一个集成了规划、工具调用、观察和验证功能的多代理框架,在农业任务中表现出色,并能更好地适应环境变化。
本文提出了一种近似算法,用于计算分段代数曲线之间的连续动态时间规整(CDTW)距离,以解决现有方法对采样率和异常值敏感的问题。
Existing pruning methods struggle to simultaneously preserve task performance and biological plausibility, particularly in functional recurrent neural networks. This work proposes and empirically validates noise-prune, a novel local pruning approach based on synaptic noise fluctuations: it retains critical connections through local sampling and rescales their weights to maintain average synaptic strength. The study demonstrates that both the connection sampling strategy and the weight rescaling mechanism are essential for performance, and revises the theoretically predicted optimal rescaling magnitude. Noise-prune significantly outperforms magnitude-only pruning strategies and matches or even exceeds the performance of non-local methods that rely on second-order information, while remaining computationally efficient and biologically plausible.
Existing reasoning-time strategies exhibit critical limitations: self-correction tends to reinforce initial biases, multi-agent collaboration (MAC) often suffers from insufficient coordination leading to collective errors, and high-accuracy verifiers require extensive human annotations. This paper proposes AdCo, the first framework to introduce a UCB-based adaptive “coopetition” mechanism—dynamically balancing cooperation and competition—into multi-agent LLM reasoning. AdCo leverages only coarse-grained verification signals to guide uncertainty-aware exploration and enhance trajectory diversity. Its methodology integrates multi-agent collaborative reasoning, iterative refinement, reasoning trajectory analysis, and knowledge diversity modeling. Evaluated on multiple mathematical reasoning benchmarks, AdCo achieves a 20% relative performance gain over state-of-the-art baselines while demonstrating robustness across varying sample sizes and configuration settings.