Institution profile

Shandong Jianzhu University

Academic institutionasia · cn
Official website
Research library29linked papers
Opportunities0open roles
Selected work

Representative Papers

Fine-grained Textual Inversion Network for Zero-Shot Composed Image Retrieval

Jul 10, 2024Annual International ACM SIGIR Conference on Research and Development in Information Retrieval

In zero-shot compositional image retrieval (ZS-CIR), single-pseudo-word mapping fails to capture fine-grained image semantics. To address this, we propose a fine-grained text inversion framework that decomposes an image into multiple pseudo-words—separately encoding subject and attribute semantics—and applies semantic regularization using BLIP-generated triplet-style captions. We further introduce multi-pseudo-word embedding modeling, template-guided cross-modal alignment, and contrastive learning to enhance compositional reasoning. Crucially, our method requires no annotated triplets. Evaluated on FashionIQ, CIRR, and CIRCO, it achieves significant improvements over state-of-the-art ZS-CIR approaches. These results demonstrate that fine-grained pseudo-word representations are essential for effective vision–language co-understanding in compositional retrieval tasks.

10 citations1 influentialRead paper

Agent-as-a-Judge

Jan 08, 2026arXiv.org

Traditional LLM-as-a-Judge approaches are limited in evaluating complex, multi-step tasks due to inherent biases, shallow reasoning, and a lack of real-world validation capabilities. This work proposes a paradigm shift toward Agent-as-a-Judge, establishing a unified, verifiable, and fine-grained evaluation framework for intelligent agents. Through a systematic review, the study integrates key technical dimensions—including planning, tool usage, multi-agent collaboration, and persistent memory—and presents the first comprehensive taxonomy of evaluation benchmarks spanning both general and domain-specific scenarios. Furthermore, it outlines a roadmap for future research, clearly identifying current challenges and promising directions in agent-based evaluation methodologies.

4 citationsRead paper

Multi-Source Multi-View Graph Domain Adaptation with Hyperbolic Residual Encoding for Cross-Site MDD Identification from rs-fMRI

Jul 31, 2026

This work addresses the challenge of identifying major depressive disorder (MDD) from cross-site resting-state fMRI data, where distribution shifts and heterogeneous functional connectivity views hinder generalization. To tackle this, we propose a multi-source unsupervised domain adaptation framework that uniquely integrates multi-view graph domain adaptation with hyperbolic representation learning. Our approach employs view-specific graph attention networks to extract features, combined with a dual-stream adaptive fusion mechanism, hyperbolic residual encoding, and Cauchy–Schwarz class alignment. It further incorporates adversarial learning, information maximization, and confidence-aware pseudo-labeling to jointly optimize cross-view consistency and cross-site alignment. Evaluated on seven unlabeled target domains, the method achieves an average accuracy of 73.60% and an AUC of 71.90%, significantly enhancing the generalization performance of cross-site MDD recognition.

0 citationsRead paper

FiRE: Enhancing MLLMs with Fine-Grained Context Learning for Complex Image Retrieval

Jul 30, 2026

This work addresses the limitations of existing approaches in complex image retrieval, which suffer from insufficient fine-grained contextual modeling and entangled optimization objectives, thereby constraining the performance of multimodal large language models. To overcome these challenges, the authors propose an automatic pipeline for constructing a fine-grained multimodal quintuple dataset and introduce a two-stage decoupled fine-tuning strategy: first enhancing contextual reasoning capabilities and subsequently refining retrieval alignment. The proposed method achieves substantial improvements over current state-of-the-art approaches across five complex image retrieval benchmarks. Notably, even with a lightweight backbone model under zero-shot settings, it attains leading performance, demonstrating the efficacy of fine-grained representation learning and staged optimization in multimodal retrieval tasks.

0 citationsRead paper

SpaceEra++: A Unified Framework Towards 3D Spatial Reasoning in Video

Jul 02, 2026

This work addresses key challenges in pre-trained vision-language models, including spatial ambiguity arising from 2D observations, scarcity of data for 3D spatial understanding, insufficient video input information, and weak reasoning constraints. To tackle these issues, the authors propose SpaceEra++, a unified framework that constructs compact yet semantically and spatially balanced scene representations through a novel ScenePick frame sampling strategy. SpaceEra++ further introduces the SpaceAlign mechanism, which enforces pairwise object constraints by jointly leveraging absolute coordinates and relative spatial relationships to enhance 3D reasoning. Combined with multi-stage training optimization and 3D-aware prompt engineering, the framework significantly outperforms strong baselines across multiple benchmarks. Ablation studies confirm the effectiveness of each component, establishing SpaceEra++ as a new paradigm for 3D vision-language research.

0 citationsRead paper
Recent publications

Latest Papers

Multi-Source Multi-View Graph Domain Adaptation with Hyperbolic Residual Encoding for Cross-Site MDD Identification from rs-fMRI

Jul 31, 2026

This work addresses the challenge of identifying major depressive disorder (MDD) from cross-site resting-state fMRI data, where distribution shifts and heterogeneous functional connectivity views hinder generalization. To tackle this, we propose a multi-source unsupervised domain adaptation framework that uniquely integrates multi-view graph domain adaptation with hyperbolic representation learning. Our approach employs view-specific graph attention networks to extract features, combined with a dual-stream adaptive fusion mechanism, hyperbolic residual encoding, and Cauchy–Schwarz class alignment. It further incorporates adversarial learning, information maximization, and confidence-aware pseudo-labeling to jointly optimize cross-view consistency and cross-site alignment. Evaluated on seven unlabeled target domains, the method achieves an average accuracy of 73.60% and an AUC of 71.90%, significantly enhancing the generalization performance of cross-site MDD recognition.

0 citationsRead paper

FiRE: Enhancing MLLMs with Fine-Grained Context Learning for Complex Image Retrieval

Jul 30, 2026

This work addresses the limitations of existing approaches in complex image retrieval, which suffer from insufficient fine-grained contextual modeling and entangled optimization objectives, thereby constraining the performance of multimodal large language models. To overcome these challenges, the authors propose an automatic pipeline for constructing a fine-grained multimodal quintuple dataset and introduce a two-stage decoupled fine-tuning strategy: first enhancing contextual reasoning capabilities and subsequently refining retrieval alignment. The proposed method achieves substantial improvements over current state-of-the-art approaches across five complex image retrieval benchmarks. Notably, even with a lightweight backbone model under zero-shot settings, it attains leading performance, demonstrating the efficacy of fine-grained representation learning and staged optimization in multimodal retrieval tasks.

0 citationsRead paper

SpaceEra++: A Unified Framework Towards 3D Spatial Reasoning in Video

Jul 02, 2026

This work addresses key challenges in pre-trained vision-language models, including spatial ambiguity arising from 2D observations, scarcity of data for 3D spatial understanding, insufficient video input information, and weak reasoning constraints. To tackle these issues, the authors propose SpaceEra++, a unified framework that constructs compact yet semantically and spatially balanced scene representations through a novel ScenePick frame sampling strategy. SpaceEra++ further introduces the SpaceAlign mechanism, which enforces pairwise object constraints by jointly leveraging absolute coordinates and relative spatial relationships to enhance 3D reasoning. Combined with multi-stage training optimization and 3D-aware prompt engineering, the framework significantly outperforms strong baselines across multiple benchmarks. Ablation studies confirm the effectiveness of each component, establishing SpaceEra++ as a new paradigm for 3D vision-language research.

0 citationsRead paper

OSGNet with MLLM Reranking @ Ego4D Episodic Memory Challenge 2026

May 20, 2026

This work addresses the challenge of temporal localization in long, untrimmed first-person videos given natural language queries or goal-step descriptions. The authors propose a two-stage reranking framework: an initial candidate generation stage using OSGNet, followed by a reranking stage that leverages a multimodal large language model (MLLM) for fine-grained semantic matching. This approach is the first to harness the powerful video-language reasoning capabilities of MLLMs specifically for the reranking phase, achieving significantly improved localization accuracy while maintaining high recall efficiency. The method secured first place in both the Natural Language Queries and GoalStep tracks of the Ego4D 2026 Challenge.

0 citationsRead paper

VISTA: Technical Report for the Ego4D Short-Term Object Interaction Anticipation at EgoVis 2026

May 20, 2026

This work addresses the problem of short-term human-object interaction prediction in egocentric videos, aiming to forecast future interacting objects’ bounding boxes, noun and verb categories, contact timestamps, and confidence scores. The authors propose a multi-task prediction framework that integrates spatiotemporal context by building upon the StillFast architecture. Leveraging a high-resolution final frame for object detection, the method innovatively injects frozen V-JEPA 2.1 temporal representations into the Faster R-CNN detection pipeline. This is achieved through feature modulation and region-of-interest (ROI)-level context fusion, enabling object-centric spatial awareness and temporal modeling. Combined with multi-head prediction and model ensembling, the approach secured first place in the EgoVis 2026 Ego4D Short-Term Interaction Prediction Challenge.

0 citationsRead paper