Institution profile

University of International Business and Economics

Academic institutionasia · cn
Official website
Research library64linked papers
Opportunities0open roles
Selected work

Representative Papers

EgoHieraLoc: A Cortically Inspired Hierarchical Segmentation-Guided Framework for Egocentric Visual Query Localization

Aug 10, 2026

This work addresses the challenge of visual query localization in first-person videos, where ambiguous object boundaries and insufficient global contextual guidance hinder performance. Inspired by the hierarchical perceptual mechanisms of the human cortex, the authors propose a unified 2D/3D visual query localization framework. The method leverages segmentation priors to extract foreground-aware query representations and employs deformable correlation filters for robust localization, further refining boundaries through multi-scale region-adaptive contextual feedback. In 3D scenes, it innovatively introduces a geometry–semantics joint confidence measure to evaluate and fuse multi-view information. To the best of our knowledge, this is the first study to integrate hierarchical perception with geometry–semantics confidence modeling for this task, achieving state-of-the-art performance on both VQL-2D and VQL-3D benchmarks.

0 citationsRead paper

Breaking Feedback-Blindness: Utility-Augmented Transformer for Sequential Decision Making

Jul 21, 2026

This work addresses the suboptimal decision-making of existing Transformer-based agents in non-stationary, partially observable environments, where reliance solely on observation similarity for attention retrieval fails to distinguish between distinct action–reward histories under identical observations. To overcome this limitation, the paper introduces the Utility-Augmented Transformer (UAT), which formally characterizes the “feedback-blind retrieval” problem and incorporates a utility-state-modulated attention mechanism. This mechanism explicitly integrates action–reward history into query, key, and value projections to guide contextual retrieval. The proposed architecture features a zero-gate degeneracy property, strictly generalizing the representational capacity of observation-only Transformers, and offers theoretical guarantees under Lipschitz continuity and finite-horizon assumptions. Empirical results demonstrate that UAT significantly outperforms current baselines across four non-stationary benchmark tasks, particularly excelling in high-noise regimes and rapid adaptation scenarios.

0 citationsRead paper
Recent publications

Latest Papers

EgoHieraLoc: A Cortically Inspired Hierarchical Segmentation-Guided Framework for Egocentric Visual Query Localization

Aug 10, 2026

This work addresses the challenge of visual query localization in first-person videos, where ambiguous object boundaries and insufficient global contextual guidance hinder performance. Inspired by the hierarchical perceptual mechanisms of the human cortex, the authors propose a unified 2D/3D visual query localization framework. The method leverages segmentation priors to extract foreground-aware query representations and employs deformable correlation filters for robust localization, further refining boundaries through multi-scale region-adaptive contextual feedback. In 3D scenes, it innovatively introduces a geometry–semantics joint confidence measure to evaluate and fuse multi-view information. To the best of our knowledge, this is the first study to integrate hierarchical perception with geometry–semantics confidence modeling for this task, achieving state-of-the-art performance on both VQL-2D and VQL-3D benchmarks.

0 citationsRead paper

Breaking Feedback-Blindness: Utility-Augmented Transformer for Sequential Decision Making

Jul 21, 2026

This work addresses the suboptimal decision-making of existing Transformer-based agents in non-stationary, partially observable environments, where reliance solely on observation similarity for attention retrieval fails to distinguish between distinct action–reward histories under identical observations. To overcome this limitation, the paper introduces the Utility-Augmented Transformer (UAT), which formally characterizes the “feedback-blind retrieval” problem and incorporates a utility-state-modulated attention mechanism. This mechanism explicitly integrates action–reward history into query, key, and value projections to guide contextual retrieval. The proposed architecture features a zero-gate degeneracy property, strictly generalizing the representational capacity of observation-only Transformers, and offers theoretical guarantees under Lipschitz continuity and finite-horizon assumptions. Empirical results demonstrate that UAT significantly outperforms current baselines across four non-stationary benchmark tasks, particularly excelling in high-noise regimes and rapid adaptation scenarios.

0 citationsRead paper