Quantum Reinforcement Learning Trading Agent for Sector Rotation in the Taiwan Stock Market
This study addresses the agent-objective mismatch in quantum reinforcement learning (QRL) for financial decision-making—characterized by high training rewards but poor out-of-sample performance—by proposing a hybrid quantum-classical RL framework for sector rotation in the Taiwan stock market. Methodologically, it adopts Proximal Policy Optimization (PPO) as the backbone, integrating LSTM and Transformer architectures with three quantum-enhanced modules: Quantum Neural Networks (QNNs), Quantum-Enhanced RWKV (QRWKV), and Quantum-Adaptive Self-Attention (QASA), alongside automated feature engineering, under NISQ-device constraints for end-to-end training. Contributions include: (1) establishing the first reproducible quantum-classical financial RL benchmark; (2) empirically demonstrating that short-horizon reward functions induce overfitting, resulting in significantly lower cumulative returns and Sharpe ratios versus classical baselines; and (3) identifying limited quantum circuit expressivity and optimization instability as key bottlenecks underlying the quantum-classical performance gap.