TradingMoE: Routing the Right Experts in Evolving Markets

📅 2026-08-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing large language models struggle to dynamically adapt to the evolving assets, scenarios, and market conditions inherent in financial trading, while conventional expert routing mechanisms fail to accurately assess each expert’s actual contribution to decision-making. To address these limitations, this work proposes TradingMoE—a trading-oriented sparse mixture-of-experts architecture that integrates lightweight residual experts into a frozen dense large language model. The approach introduces a novel market-context-aware query-key-value router and a market-evolution-sensitive sparse expert updating strategy, thereby overcoming the constraints of static routing. Empirical results demonstrate that TradingMoE outperforms 22 baseline models on both stock and cryptocurrency markets, achieving cumulative return improvements of 30.89% and 30.7%, respectively, and maintains consistent superiority in forward-only rolling paper trading simulations.
📝 Abstract
Large language models (LLMs) have shown strong potential for financial analysis and trading, but direct trading remains challenging because the predictive capabilities required can vary across assets, decision fields, and market conditions. Existing LLM-based trading systems either coordinate human-defined external experts or adopt conventional internal Mixture-of-Experts (MoE) routers that do not directly evaluate how individual experts contribute to trading decisions. Moreover, these routers receive no direct signal indicating when an inactive expert has become more suitable as market conditions change. We find that native router scores poorly reflect how much individual experts improve trading decisions, frequently leaving better alternatives unselected. We further reveal that token-specific expert usefulness exhibits a compact low-dimensional structure. Based on these findings, we propose TradingMoE, a trading-oriented sparse MoE that augments a frozen dense LLM with lightweight residual experts. We introduce a Query-Key router that represents the expertise required by each token under the current market context as a low-dimensional query and matches it with learnable expert keys. We further propose a sparse expert selection update mechanism that samples a few inactive experts during training and estimates whether they should replace the weakest expert in the current Top-k route. This mechanism enables the router to update expert selection as market conditions change while preserving sparse computation. Experiments against 22 baselines on stock and cryptocurrency markets show that TradingMoE improves cumulative return over the best-performing baselines by 30.89% and 30.7%, respectively. Rolling paper-trading experiments further demonstrate that its advantage persists under forward-only deployment.
Problem

Research questions and friction points this paper is trying to address.

trading
Mixture-of-Experts
expert routing
market dynamics
large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

TradingMoE
Query-Key router
sparse Mixture-of-Experts
expert selection update
low-dimensional expertise structure
🔎 Similar Papers
No similar papers found.