Training Skills Like Parameters via Self-Supervised Semantic Diffusion
This work addresses the limitations of existing closed-source large language models in efficiently acquiring, updating, and reusing domain-specific skills, which often rely on costly human annotations or unreliable model-based evaluations. The authors propose an unsupervised self-evolving agent framework that, for the first time, adapts the destroy-and-reconstruct mechanism from diffusion models to skill learning. By contrasting agent-reconstructed text with original high-quality human-written text, the framework generates self-supervised signals to refine an external skill library—comprising skills that are readable, transferable, and composable—without updating the model’s internal weights. Integrating self-supervised contrastive learning with a semantic diffusion mechanism, the approach enables autonomous skill distillation and continuous iteration. Evaluated on short-form screenplay generation, it significantly improves output quality, demonstrating strong scalability and generalization in autonomously learning complex human artifacts.