Mirror Learning
This work addresses the challenge that existing behavior cloning methods rely on first-person aligned data and struggle to learn effective policies from third-person passive observations. The authors propose a Mirror Learning framework that, for the first time, integrates viewpoint transformation with inverse dynamics modeling. By fine-tuning a video diffusion model to translate third-person observations into first-person perspectives and employing an inverse dynamics model to infer action trajectories, the method generates pseudo-first-person expert demonstrations from purely observational videos. This approach constructs a generative world model capable of training high-performance policies using only mirrored data, substantially reducing reliance on teleoperated demonstrations. When combined with first-person behavior cloning, the framework further enhances downstream policy performance.