Causal Object-Centric Models for Planning with Monte Carlo Tree Search
This work addresses the challenges of low sample efficiency and insufficient decision focus in visual reinforcement learning by proposing an object-centric planning framework. Operating within a slot-structured latent space derived from a frozen object encoder, the method integrates a Transformer-based world model with Monte Carlo Tree Search (MCTS), augmented by an action-slot fusion mechanism to accurately predict object-level state transitions. Furthermore, an object-causal attention mechanism dynamically guides the policy and value networks to attend to task-relevant entities. By explicitly incorporating object-level inductive biases, the approach achieves substantial performance gains over both object-centric and monolithic baselines across eight benchmark tasks in Object-Centric Visual RL, ManiSkill, RoboSuite, and VizDoom, demonstrating notably higher average normalized scores especially during early training stages.