Aspire: Can Models Self-Evolve from Vague Goals?

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究提出ASPIRE框架,让模型从模糊目标自我进化,通过自选数据和方法学习,解决现有模型依赖明确任务指标的问题。
📝 Abstract
Many important forms of human learning begin with a vague goal, such as "become a better physicist" or "improve at research." Learners must interpret the goal, identify capability gaps, decide how to learn, and determine whether they have actually improved. In contrast, existing work on LLM self-evolution typically begins with tasks and evaluation metrics specified by humans, reducing self-evolution to optimizing an explicit objective rather than deciding what and how to learn. We introduce ASPIRE, a benchmark for vague-goal-driven self-evolution. ASPIRE provides only a natural-language capability goal while downstream evaluation tasks remain hidden. The agent must operationalize the goal by choosing data and update methods, constructing training and validation signals, and deciding when to evaluate. ASPIRE supports both model-weight and agent-harness evolution in a unified interactive environment and evaluates the resulting systems on a hidden, expert-authored set of 520 items spanning six goals. Our experiments show that vague goals redirect search effort toward goal interpretation. Current agents routinely complete training and harness-editing loops, but weight-level gains remain sparse and unstable, and the strongest evolved harness remains below the engineered Qwen-Agent reference. Agents often train on mismatched data and trust narrow self-evaluations, so local gains fail to transfer to hidden evaluation and continued search and training can erase earlier improvements.
Problem

Research questions and friction points this paper is trying to address.

vague goals
self-evolution
LLM
capability gaps
natural-language capability goal
Innovation

Methods, ideas, or system contributions that make the work stand out.

vague-goal-driven self-evolution
natural-language capability goal
unified interactive environment
model-weight and agent-harness evolution
Y
Yuhao Wu
ByteDance Seed
J
Jingyuan Zhang
Singapore University of Technology and Design
J
Jiajun Shi
M-A-P
Y
Yuxuan Zhang
TokenWave.AI
X
Xinping Lei
ByteDance Seed
Junting Zhou
Junting Zhou
Peking University
Large Language ModelAI for ScienceBioinformatics
Z
Zexuan Wang
M-A-P
Y
Yuchen Wu
TokenWave.AI
Huan Zhou
Huan Zhou
Northwestern Polytechnical University
Mobile Edge ComputingFederated LearningMobile Social NetworksVANETsData Offloading
D
Duo Wang
Singapore University of Technology and Design
Y
Yinzhu Piao
M-A-P
Y
Yongchang Peng
TokenWave.AI
Y
Yunfeng Shi
ByteDance Seed
J
Jin Chen
Singapore University of Technology and Design
Z
Zuo Wang
M-A-P
J
Jinkai Liu
TokenWave.AI
J
Jiaheng Liu
ByteDance Seed
Wenxuan Zhang
Wenxuan Zhang
Singapore University of Technology and Design
Natural Language ProcessingLarge Language ModelsMultilingual NLP
S
Shen Yan
M-A-P
W
Wenhao Huang
TokenWave.AI
G
Ge Zhang
ByteDance Seed