PerspAct: Enhancing LLM Situated Collaboration Skills through Perspective Taking and Active Vision
Current large language models (LLMs) and multimodal models exhibit limited perspective-taking capabilities in multi-agent collaboration, hindering accurate modeling of subjective agent perceptions and multi-observer environments. To address this, we propose PerspAct—a novel method that integrates active visual exploration with the ReAct reasoning framework for the first time. PerspAct explicitly samples and models diverse agent-centric perspectives, enabling dynamic comprehension of hierarchical perspective complexity in an extended Director task. Built upon multimodal LLMs, it leverages prompt engineering and explicit state representation. We systematically evaluate PerspAct across seven progressively complex scenarios. Experiments demonstrate significant improvements in both coreference resolution and collaborative task accuracy, validating the efficacy of jointly modeling active perception and perspective understanding. Our work establishes a new paradigm for situational awareness in multi-agent settings.