Self-Predictive Dynamics for Generalization of Vision-based Reinforcement Learning
Vision-based reinforcement learning exhibits poor generalization under unseen image corruptions (e.g., shadows, cloud cover, illumination shifts). To address this, we propose Self-Predictive Dynamics (SPD), the first method to jointly integrate bidirectional (forward and inverse) dynamics prediction with weak/strong dual-path data augmentation for task-agnostic, robust representation learning. SPD employs contrastive self-supervised modeling to explicitly disentangle task-relevant dynamics from corruption-invariant features. Evaluated on MuJoCo vision-based control and CARLA autonomous driving benchmarks, SPD significantly improves policy generalization across corrupted environments—achieving a 23.6% performance gain over state-of-the-art methods under unseen corruptions. The implementation is publicly available.