Gaussian Processes for Modelling Spatial Fields with Robot Swarms
本文提出了一种位置未知的高斯过程回归方法,用于在没有外部定位系统的情况下通过机器人集群建模空间场,如水流温度、风速或地形高度。
本文提出了一种位置未知的高斯过程回归方法,用于在没有外部定位系统的情况下通过机器人集群建模空间场,如水流温度、风速或地形高度。
为了解决视频中错误检测的问题,特别是对于未见过的动作,提出了一种基于视频-语言模型的后训练技术,并通过特定奖励函数提高模型识别指令与视频之间差异的能力。
This work addresses the absence of a systematic evaluation benchmark for general-purpose multimodal large language models (MLLMs) in structured visual grounding tasks. It proposes the first promptable visual grounding evaluation framework tailored for such models, encompassing four task categories: object detection, referring expression comprehension, instance-level localization, and video grounding. The framework employs a unified input template, standardized bounding box output format, and a consistent cross-task evaluation protocol. Notably, it explicitly emphasizes the model’s adherence to prescribed output formats—a previously overlooked yet critical capability. Systematic evaluations reveal that prevailing MLLMs are highly sensitive to output formatting and exhibit limited generalization across tasks, thereby highlighting key directions for future model improvement.
Existing attention models struggle to explicitly capture the temporal dynamics of human gaze in dynamic scenes, often relying on saliency maps or scanpaths for implicit modeling. This work proposes an autoregressive dynamical system that formulates gaze trajectories as a generative process jointly driven by gaze history and the evolving environment, centered around a novel gaze-centric heterogeneous graph structure. To this end, we introduce the Heterogeneous Graph Transformer (ART) and the Object Density Network (ODN), enabling, for the first time, unified modeling of natural gaze trajectories, scanpaths, and saliency maps directly from unfiltered raw gaze data. Evaluated on the newly released Focus100 driving gaze dataset, our model, trained end-to-end, generates more naturalistic gaze trajectories while significantly improving the dynamic fidelity of scanpaths and the accuracy of saliency predictions, thereby achieving precise modeling of the temporal characteristics of human attention in dynamic environments.
Existing vision-language models often rely on additional training or external modules for personalization, which limits their generalization, scalability, and deployment efficiency. This work proposes a lightweight, fine-tuning-free approach that leverages the model’s intrinsic attention mechanisms to automatically extract visual tokens representing target concepts as “concept memories.” During inference, these memories enable efficient personalized responses through embedding guidance. The method unifies support for single-concept, multi-concept, and video-based scenarios, consistently outperforming current state-of-the-art approaches across diverse settings. Notably, it achieves high performance with minimal computational overhead, demonstrating strong generality and practical utility.
本文提出了一种位置未知的高斯过程回归方法,用于在没有外部定位系统的情况下通过机器人集群建模空间场,如水流温度、风速或地形高度。
为了解决视频中错误检测的问题,特别是对于未见过的动作,提出了一种基于视频-语言模型的后训练技术,并通过特定奖励函数提高模型识别指令与视频之间差异的能力。
This work addresses the absence of a systematic evaluation benchmark for general-purpose multimodal large language models (MLLMs) in structured visual grounding tasks. It proposes the first promptable visual grounding evaluation framework tailored for such models, encompassing four task categories: object detection, referring expression comprehension, instance-level localization, and video grounding. The framework employs a unified input template, standardized bounding box output format, and a consistent cross-task evaluation protocol. Notably, it explicitly emphasizes the model’s adherence to prescribed output formats—a previously overlooked yet critical capability. Systematic evaluations reveal that prevailing MLLMs are highly sensitive to output formatting and exhibit limited generalization across tasks, thereby highlighting key directions for future model improvement.
Existing attention models struggle to explicitly capture the temporal dynamics of human gaze in dynamic scenes, often relying on saliency maps or scanpaths for implicit modeling. This work proposes an autoregressive dynamical system that formulates gaze trajectories as a generative process jointly driven by gaze history and the evolving environment, centered around a novel gaze-centric heterogeneous graph structure. To this end, we introduce the Heterogeneous Graph Transformer (ART) and the Object Density Network (ODN), enabling, for the first time, unified modeling of natural gaze trajectories, scanpaths, and saliency maps directly from unfiltered raw gaze data. Evaluated on the newly released Focus100 driving gaze dataset, our model, trained end-to-end, generates more naturalistic gaze trajectories while significantly improving the dynamic fidelity of scanpaths and the accuracy of saliency predictions, thereby achieving precise modeling of the temporal characteristics of human attention in dynamic environments.
Existing vision-language models often rely on additional training or external modules for personalization, which limits their generalization, scalability, and deployment efficiency. This work proposes a lightweight, fine-tuning-free approach that leverages the model’s intrinsic attention mechanisms to automatically extract visual tokens representing target concepts as “concept memories.” During inference, these memories enable efficient personalized responses through embedding guidance. The method unifies support for single-concept, multi-concept, and video-based scenarios, consistently outperforming current state-of-the-art approaches across diverse settings. Notably, it achieves high performance with minimal computational overhead, demonstrating strong generality and practical utility.