OpenThoughts-Agent: Data Recipes for Agentic Models
This work addresses the lack of open, generalizable methodologies for constructing training data for intelligent agents—a key limitation hindering their generalization across diverse tasks. The study presents the first systematic investigation into agent training data construction, introducing an open-source and scalable data recipe. Through multi-source task sampling, diversity optimization, and controlled ablation studies, the authors rigorously analyze how task provenance and data composition influence model performance. A 100K-sample training set built using this approach achieves an average accuracy of 44.8% across seven agent benchmarks, outperforming the strongest existing open-source model by 3.9 percentage points. The method consistently maintains superior performance across varying data scales, demonstrating strong generalization and practical utility.