Institution profile

Shandong University

Academic institutionasia · cn
Official website
Research library1,097linked papers
Opportunities0open roles
Selected work

Representative Papers

Continuous Input Embedding Size Search For Recommender Systems

Apr 07, 2023Annual International ACM SIGIR Conference on Research and Development in Information Retrieval

To address memory inefficiency caused by fixed high-dimensional embeddings in recommender systems, this paper proposes a memory-constrained continuous embedding dimension optimization framework. Unlike conventional approaches that employ uniform high-dimensional embeddings or existing reinforcement learning (RL)-based methods limited to discrete dimension selection, our work introduces the first continuous-space embedding dimension search paradigm. We design a stochastic walk-driven exploration strategy to efficiently navigate the continuous dimension space, enabling joint optimization of recommendation accuracy and memory efficiency. The method is model-agnostic and plug-and-play. Extensive experiments on two real-world datasets and three state-of-the-art recommendation models demonstrate that our approach achieves superior performance across multiple memory budgets, consistently outperforming discrete-search baselines and establishing new state-of-the-art results.

21 citations1 influentialRead paper

Dynamic Multimodal Fusion via Meta-Learning Towards Micro-Video Recommendation

Aug 30, 2023ACM Trans. Inf. Syst.

To address the limitation of static multimodal fusion in middle-school micro-video recommendation—its inability to capture inter-video modality relationship discrepancies—this paper proposes MetaMMF, a meta-learning-based dynamic multimodal fusion framework. Methodologically, MetaMMF treats multimodal fusion for each video as an individual meta-task and employs meta-learning to generate video-specific fusion functions; it further adopts CP tensor decomposition to enhance parameter efficiency and training stability. While implicitly incorporating graph neural network principles (e.g., akin to MMGCN), MetaMMF avoids explicit graph construction. Extensive experiments on three benchmark datasets demonstrate that MetaMMF consistently outperforms state-of-the-art models—including MMGCN, LATTICE, and InvRL—achieving superior recommendation accuracy and computational efficiency. The source code is publicly released, empirically validating the dual advantages of dynamic fusion in both performance and efficiency.

15 citationsRead paper

Fine-grained Textual Inversion Network for Zero-Shot Composed Image Retrieval

Jul 10, 2024Annual International ACM SIGIR Conference on Research and Development in Information Retrieval

In zero-shot compositional image retrieval (ZS-CIR), single-pseudo-word mapping fails to capture fine-grained image semantics. To address this, we propose a fine-grained text inversion framework that decomposes an image into multiple pseudo-words—separately encoding subject and attribute semantics—and applies semantic regularization using BLIP-generated triplet-style captions. We further introduce multi-pseudo-word embedding modeling, template-guided cross-modal alignment, and contrastive learning to enhance compositional reasoning. Crucially, our method requires no annotated triplets. Evaluated on FashionIQ, CIRR, and CIRCO, it achieves significant improvements over state-of-the-art ZS-CIR approaches. These results demonstrate that fine-grained pseudo-word representations are essential for effective vision–language co-understanding in compositional retrieval tasks.

10 citations1 influentialRead paper

Towards Urban General Intelligence: A Review and Outlook of Urban Foundation Models

Jan 30, 2024arXiv.org

Current Urban Foundation Models (UFMs) lack formal definitions, systematic taxonomies, and a unifying framework, hindering the advancement of Urban General Intelligence (UGI). Method: This work formally defines UFMs for the first time and establishes the first taxonomy driven by multimodal urban data. It proposes a general UGI-oriented framework supporting cross-modal understanding, cross-task transfer, and cross-regional co-evolution—integrating data-center paradigms, abstract foundation model architectures, cross-domain alignment techniques, and open-source collaborative governance mechanisms. Contributions/Results: We introduce a unified conceptual paradigm and research landscape; launch Awesome-Urban-Foundation-Models—a continuously updated, authoritative resource repository; and deliver reusable methodologies and technical roadmaps. Collectively, these advances enable a paradigm shift in smart cities—from narrow, task-specific AI toward generalized, collaborative, and self-evolving urban intelligence.

7 citationsRead paper

HUD: Hierarchical Uncertainty-Aware Disambiguation Network for Composed Video Retrieval

Oct 27, 2025Proceedings of the 33rd ACM International Conference on Multimedia

To address referential ambiguity and insufficient fine-grained semantic attention in Compositional Video Retrieval (CVR), caused by cross-modal information density disparities between reference videos and modification texts, this paper proposes the Hierarchical Uncertainty-aware Disambiguation network (HUD). HUD is the first to explicitly leverage inter-modal information density differences to jointly perform holistic pronoun disambiguation, atomic-level uncertainty modeling, and progressive “holistic → atomic” alignment—integrating cross-modal interaction with fine-grained semantic alignment. Evaluated on multiple CVR and Compositional Image Retrieval (CIR) benchmarks, HUD achieves state-of-the-art performance with strong generalization capability. Its core contributions are: (1) uncovering and formally modeling how modality-specific information density differences impede multimodal understanding; and (2) introducing a hierarchical, uncertainty-driven paradigm for disambiguation and alignment that bridges coarse- and fine-grained semantic representations across modalities.

3 citationsRead paper
Recent publications

Latest Papers