Hypothesis-Driven Skill Optimization for LLM Agents
This work addresses the vulnerability of external skill updating to sparse or noisy execution trajectories, which can lead to the entrenchment of ineffective or even detrimental policies. To mitigate this issue, the authors propose a training-free skill optimization framework that leverages a falsifiable hypothesis-driven mechanism. By integrating controlled experiments, behavioral discrepancy analysis, and progressive skill disclosure, the method enables auditable and noise-resilient skill curation and execution at frozen model inference endpoints. Evaluated on ALFWorld, the approach yields substantial performance gains: average success rates improve by 6.9 and 4.0 percentage points for Qwen3-8B and Qwen3.6-27B, respectively. Notably, it maintains a +7.1-point advantage even under 20% erroneous feedback and demonstrates strong cross-run and cross-model transferability.