OGR-MARL: Option-Guided Residual Multi-Agent Reinforcement Learning for Heterogeneous USV Cooperative Pursuit in Constrained Port Waterways
This work addresses the challenge of cooperative pursuit by heterogeneous unmanned surface vehicles (USVs) in constrained port waterways, where navigation, traffic, and role-based constraints must be respected. To this end, the authors propose the OGR-MARL framework, which integrates shared goal beliefs, a role-conditioned options mechanism, adaptive rule-based penalties, and residual policy learning. This enables multi-agent systems to refine actions while adhering to prior domain rules, rather than exploring from scratch. Notably, the framework decouples the options mechanism and residual learning from specific multi-agent reinforcement learning algorithms—such as MADDPG, MATD3, MAPPO, and MASAC—yielding a general architecture capable of zero-shot transfer to real-world port maps. Evaluated in the simulated Shin-no-Mon port environment, OGR-MASAC achieves a 75.0% capture rate, substantially outperforming baselines, and successfully generalizes without retraining to realistic scenarios derived from QGIS/AIS data, demonstrating both effective coordination and strict rule compliance.