π€ AI Summary
This study addresses the subjectivity and lack of end-to-end feedback in quantitative factor mining by proposing a Large Language Model-driven automated factor generation framework. We introduce a novel trajectory-level evolutionary mechanism combined with self-iterative redundancy-aware ensemble, integrating multi-source information summarization and evolutionary operator search to achieve closed-loop optimization from hypothesis formulation to backtesting. Experiments on the CSI300 index demonstrate superior performance with an annualized return rate of 8.28%, an information ratio of 1.29, and an information coefficient of 0.0454. Furthermore, zero-shot evaluation on the CSI500 validates strong cross-market transferability. These results confirm that the proposed approach significantly enhances both the robustness and generalizability of automated factor discovery processes.
π Abstract
With the rapid rise of large language models, LLM-driven quantitative factor mining has become an increasingly active research area. However, existing methods still suffer from subjective direction design, limited integration of up-to-date multi-source information, semantic drift, factor redundancy, and the absence of an end-to-end feedback loop from factor discovery to portfolio backtesting. To address these limitations, we propose AlphaSeek, an end-to-end factor mining framework for quantitative investment that integrates automated direction discovery, trajectory-level factor evolution mining and self-iterative portfolio optimization. AlphaSeek first collects and summarizes multi-source financial information to identify promising mining directions. It then performs trajectory-level factor mining by extending the optimization unit from a single factor expression to a complete research trajectory covering hypothesis generation, factor construction, validation, backtesting, and feedback. Based on this design, we introduce evolution operators - parallel direction expansion, mutation and crossover - to improve search diversity, refinement quality and factor robustness. Finally, AlphaSeek constructs a self-iterative factor portfolio, allowing newly discovered factors to interact with an existing state-of-the-art(SOTA) factor library under redundancy-aware constraints. Experiments on CSI300 show that AlphaSeek achieves the strongest overall strategy-level performance on CSI300 with ARR of 8.28%, IR of 1.29 and MDD of 6.28%, while remaining competitive on factor predictive metrics with IC of 0.0454, while factors mined on CSI300 also achieve strong time-series return performance on CSI500 than other models, suggesting promising cross-market transferability under a zero-shot setting.