STRETCH the Boundaries: A Unified Self-Taught Framework for Progressive LLM Evolution

📅 2026-09-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决大语言模型自我提升时能力停滞问题,提出STRETCH框架,通过动态调整挑战难度与模型能力匹配,促进模型持续进化。
📝 Abstract
Large language models (LLMs) often suffer from capability stagnation in self-improvement training because fixed difficulty levels fail to adapt to their evolving proficiency. To address this issue, we propose STRETCH (Self-Taught Reasoning Evolution via Targeted CHallenge), a unified framework inspired by cognitive scaffolding theory. STRETCH introduces a dynamic Stretch Zone mechanism that continuously aligns question difficulty with the model's solving capability. Within a single parameter space, the model alternates between a Scaffolder that generates adaptive, boundary-pushing challenges and a Learner that that optimizes its solving trajectories through reinforcement learning. This dual-loop co-evolution effectively stabilizes training, mitigates reward hacking and promote progressive reasoning growth. Experiments on both negotiation and operation research benchmarks demonstrate that STRETCH consistently outperforms strong prompting and domain-specific baselines. Further scaffolder configuration analysis shows that dynamic difficulty alignment is critical for sustained capability improvement and synchronized reasoning evolution.
Problem

Research questions and friction points this paper is trying to address.

large language models
capability stagnation
self-improvement training
fixed difficulty levels
Innovation

Methods, ideas, or system contributions that make the work stand out.

dynamic Stretch Zone
Scaffolder and Learner co-evolution
reinforcement learning
🔎 Similar Papers
No similar papers found.