Institution profile

TME

Industry research
Research library6linked papers
Opportunities0open roles
Selected work

Representative Papers

Post-Training VLMs for Video Mistake Detection

Aug 28, 2026

为了解决视频中错误检测的问题,特别是对于未见过的动作,提出了一种基于视频-语言模型的后训练技术,并通过特定奖励函数提高模型识别指令与视频之间差异的能力。

0 citationsRead paper

FindIt: A Format-Informed Visual Detection Benchmark for Generalist Multimodal LLMs

Jun 02, 2026

This work addresses the absence of a systematic evaluation benchmark for general-purpose multimodal large language models (MLLMs) in structured visual grounding tasks. It proposes the first promptable visual grounding evaluation framework tailored for such models, encompassing four task categories: object detection, referring expression comprehension, instance-level localization, and video grounding. The framework employs a unified input template, standardized bounding box output format, and a consistent cross-task evaluation protocol. Notably, it explicitly emphasizes the model’s adherence to prescribed output formats—a previously overlooked yet critical capability. Systematic evaluations reveal that prevailing MLLMs are highly sensitive to output formatting and exhibit limited generalization across tasks, thereby highlighting key directions for future model improvement.

0 citationsRead paper

Beyond Scanpaths: Graph-Based Gaze Simulation in Dynamic Scenes

Mar 30, 2026

Existing attention models struggle to explicitly capture the temporal dynamics of human gaze in dynamic scenes, often relying on saliency maps or scanpaths for implicit modeling. This work proposes an autoregressive dynamical system that formulates gaze trajectories as a generative process jointly driven by gaze history and the evolving environment, centered around a novel gaze-centric heterogeneous graph structure. To this end, we introduce the Heterogeneous Graph Transformer (ART) and the Object Density Network (ODN), enabling, for the first time, unified modeling of natural gaze trajectories, scanpaths, and saliency maps directly from unfiltered raw gaze data. Evaluated on the newly released Focus100 driving gaze dataset, our model, trained end-to-end, generates more naturalistic gaze trajectories while significantly improving the dynamic fidelity of scanpaths and the accuracy of saliency predictions, thereby achieving precise modeling of the temporal characteristics of human attention in dynamic environments.

0 citationsRead paper

Ego: Embedding-Guided Personalization of Vision-Language Models

Mar 10, 2026

Existing vision-language models often rely on additional training or external modules for personalization, which limits their generalization, scalability, and deployment efficiency. This work proposes a lightweight, fine-tuning-free approach that leverages the model’s intrinsic attention mechanisms to automatically extract visual tokens representing target concepts as “concept memories.” During inference, these memories enable efficient personalized responses through embedding guidance. The method unifies support for single-concept, multi-concept, and video-based scenarios, consistently outperforming current state-of-the-art approaches across diverse settings. Notably, it achieves high performance with minimal computational overhead, demonstrating strong generality and practical utility.

0 citationsRead paper
Recent publications

Latest Papers

Post-Training VLMs for Video Mistake Detection

Aug 28, 2026

为了解决视频中错误检测的问题,特别是对于未见过的动作,提出了一种基于视频-语言模型的后训练技术,并通过特定奖励函数提高模型识别指令与视频之间差异的能力。

0 citationsRead paper

FindIt: A Format-Informed Visual Detection Benchmark for Generalist Multimodal LLMs

Jun 02, 2026

This work addresses the absence of a systematic evaluation benchmark for general-purpose multimodal large language models (MLLMs) in structured visual grounding tasks. It proposes the first promptable visual grounding evaluation framework tailored for such models, encompassing four task categories: object detection, referring expression comprehension, instance-level localization, and video grounding. The framework employs a unified input template, standardized bounding box output format, and a consistent cross-task evaluation protocol. Notably, it explicitly emphasizes the model’s adherence to prescribed output formats—a previously overlooked yet critical capability. Systematic evaluations reveal that prevailing MLLMs are highly sensitive to output formatting and exhibit limited generalization across tasks, thereby highlighting key directions for future model improvement.

0 citationsRead paper

Beyond Scanpaths: Graph-Based Gaze Simulation in Dynamic Scenes

Mar 30, 2026

Existing attention models struggle to explicitly capture the temporal dynamics of human gaze in dynamic scenes, often relying on saliency maps or scanpaths for implicit modeling. This work proposes an autoregressive dynamical system that formulates gaze trajectories as a generative process jointly driven by gaze history and the evolving environment, centered around a novel gaze-centric heterogeneous graph structure. To this end, we introduce the Heterogeneous Graph Transformer (ART) and the Object Density Network (ODN), enabling, for the first time, unified modeling of natural gaze trajectories, scanpaths, and saliency maps directly from unfiltered raw gaze data. Evaluated on the newly released Focus100 driving gaze dataset, our model, trained end-to-end, generates more naturalistic gaze trajectories while significantly improving the dynamic fidelity of scanpaths and the accuracy of saliency predictions, thereby achieving precise modeling of the temporal characteristics of human attention in dynamic environments.

0 citationsRead paper

Ego: Embedding-Guided Personalization of Vision-Language Models

Mar 10, 2026

Existing vision-language models often rely on additional training or external modules for personalization, which limits their generalization, scalability, and deployment efficiency. This work proposes a lightweight, fine-tuning-free approach that leverages the model’s intrinsic attention mechanisms to automatically extract visual tokens representing target concepts as “concept memories.” During inference, these memories enable efficient personalized responses through embedding guidance. The method unifies support for single-concept, multi-concept, and video-based scenarios, consistently outperforming current state-of-the-art approaches across diverse settings. Notably, it achieves high performance with minimal computational overhead, demonstrating strong generality and practical utility.

0 citationsRead paper