Institution profile

CogAI Lab

Research institution
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

Causal Object-Centric Models for Planning with Monte Carlo Tree Search

Jun 12, 2026

This work addresses the challenges of low sample efficiency and insufficient decision focus in visual reinforcement learning by proposing an object-centric planning framework. Operating within a slot-structured latent space derived from a frozen object encoder, the method integrates a Transformer-based world model with Monte Carlo Tree Search (MCTS), augmented by an action-slot fusion mechanism to accurately predict object-level state transitions. Furthermore, an object-causal attention mechanism dynamically guides the policy and value networks to attend to task-relevant entities. By explicitly incorporating object-level inductive biases, the approach achieves substantial performance gains over both object-centric and monolithic baselines across eight benchmark tasks in Object-Centric Visual RL, ManiSkill, RoboSuite, and VizDoom, demonstrating notably higher average normalized scores especially during early training stages.

0 citationsRead paper

$μ$VLA: On Recurrent Memory for Partially Observable Manipulation in VLA Models

Jun 10, 2026

This work addresses the failure of vision-language-action (VLA) models in partially observable environments due to the absence of historical memory. To remedy this, the authors propose an extremely lightweight recurrent memory mechanism that introduces only a small set of learnable memory tokens into a pretrained VLA backbone. These tokens are updated across timesteps via self-attention and enable end-to-end training without any architectural modifications or auxiliary losses. This approach facilitates, for the first time, a controlled and isolated study of recurrence in VLA systems. Experiments demonstrate substantial performance gains: on MIKASA-Robo, average success rates on trained tasks improve from 0.42 to 0.84, while zero-shot task performance rises to 0.23 from a baseline of 0.07. Moreover, the method maintains a high success rate of 96.2% on fully observable LIBERO tasks.

0 citationsRead paper

Learning to Communicate Locally for Large-Scale Multi-Agent Pathfinding

May 08, 2026

This work addresses the suboptimality and scalability limitations in large-scale multi-agent path finding (MAPF) stemming from insufficient coordination by proposing LC-MAPF, a novel framework that formulates MAPF as a decentralized partially observable Markov decision process with a learnable local communication mechanism. Integrating imitation learning and reinforcement learning, LC-MAPF employs graph neural networks to iteratively aggregate neighborhood information through multiple rounds of local message passing and generate coordinated actions. Experimental results demonstrate that LC-MAPF significantly outperforms existing learning-based solvers across diverse unseen scenarios, achieving higher success rates and superior path quality while maintaining strong scalability.

0 citationsRead paper
Recent publications

Latest Papers

Causal Object-Centric Models for Planning with Monte Carlo Tree Search

Jun 12, 2026

This work addresses the challenges of low sample efficiency and insufficient decision focus in visual reinforcement learning by proposing an object-centric planning framework. Operating within a slot-structured latent space derived from a frozen object encoder, the method integrates a Transformer-based world model with Monte Carlo Tree Search (MCTS), augmented by an action-slot fusion mechanism to accurately predict object-level state transitions. Furthermore, an object-causal attention mechanism dynamically guides the policy and value networks to attend to task-relevant entities. By explicitly incorporating object-level inductive biases, the approach achieves substantial performance gains over both object-centric and monolithic baselines across eight benchmark tasks in Object-Centric Visual RL, ManiSkill, RoboSuite, and VizDoom, demonstrating notably higher average normalized scores especially during early training stages.

0 citationsRead paper

$μ$VLA: On Recurrent Memory for Partially Observable Manipulation in VLA Models

Jun 10, 2026

This work addresses the failure of vision-language-action (VLA) models in partially observable environments due to the absence of historical memory. To remedy this, the authors propose an extremely lightweight recurrent memory mechanism that introduces only a small set of learnable memory tokens into a pretrained VLA backbone. These tokens are updated across timesteps via self-attention and enable end-to-end training without any architectural modifications or auxiliary losses. This approach facilitates, for the first time, a controlled and isolated study of recurrence in VLA systems. Experiments demonstrate substantial performance gains: on MIKASA-Robo, average success rates on trained tasks improve from 0.42 to 0.84, while zero-shot task performance rises to 0.23 from a baseline of 0.07. Moreover, the method maintains a high success rate of 96.2% on fully observable LIBERO tasks.

0 citationsRead paper

Learning to Communicate Locally for Large-Scale Multi-Agent Pathfinding

May 08, 2026

This work addresses the suboptimality and scalability limitations in large-scale multi-agent path finding (MAPF) stemming from insufficient coordination by proposing LC-MAPF, a novel framework that formulates MAPF as a decentralized partially observable Markov decision process with a learnable local communication mechanism. Integrating imitation learning and reinforcement learning, LC-MAPF employs graph neural networks to iteratively aggregate neighborhood information through multiple rounds of local message passing and generate coordinated actions. Experimental results demonstrate that LC-MAPF significantly outperforms existing learning-based solvers across diverse unseen scenarios, achieving higher success rates and superior path quality while maintaining strong scalability.

0 citationsRead paper