Institution profile

Pazhou Lab

Academic institutionasia · cn
Official website
Research library97linked papers
Opportunities0open roles
Selected work

Representative Papers

Learning Pyramid-structured Long-range Dependencies for 3D Human Pose Estimation

Jun 03, 2025IEEE transactions on multimedia

Modeling long-range dependencies in 3D human pose estimation remains challenging due to noise susceptibility and high model complexity. To address this, we propose the Pyramid Graph Attention (PGA) module and a lightweight Multi-Scale Graph Transformer (PGFormer). Our core contribution is the first formulation of human anatomical substructures—joints, limbs, and torso—as a pyramid-shaped, cross-scale graph, coupled with a pooling-augmented self-attention mechanism that preserves structural priors while enabling multi-granularity feature interaction. By integrating graph convolutional operations with multi-scale feature fusion, our method effectively suppresses redundancy in deep networks. Evaluated on Human3.6M and MPI-INF-3DHP, it achieves state-of-the-art accuracy (MPJPE: 41.2 mm and 89.7 mm, respectively) with a 23% reduction in parameter count, demonstrating both the efficacy and efficiency of cross-scale graph modeling for 3D pose estimation.

2 citationsRead paper

Visual Environment-Interactive Planning for Embodied Complex-Question Answering

Apr 01, 2025IEEE transactions on circuits and systems for video technology (Print)

Embodied robots face significant challenges in answering structured, abstract natural-language questions within complex visual environments. Method: We propose an environment-driven, multi-step interactive planning framework. It constructs a hierarchical visual scene graph to parse question semantics, introduces a chain-based essential question representation for intent modeling, and integrates external rules with real-time visual feedback for closed-loop sequential decision-making. Crucially, we pioneer a structured semantic space that enables iterative vision–language interaction, eliminating reliance on single-step large language models. Contribution/Results: Evaluated on a newly constructed complex embodied question-answering dataset, our approach substantially improves planning interpretability, adaptability, and robustness. Empirical results demonstrate its feasibility and practicality in realistic embodied settings.

1 citationsRead paper

Share First, Route What Remains: A Unified Framework for Token-Adaptive MoE Computation

Aug 10, 2026

This work addresses the disconnect in existing Mixture-of-Experts (MoE) models between shared computation and dynamic routing, which overlooks the interdependence between reusable computation and residual expert requirements. The paper proposes UniF-MoE, a unified framework introducing a novel “shared-first, routed-later” mechanism: it first processes common features through a shared general-purpose module and then dynamically activates residual experts based on a shared-demand score and complementarity. Key innovations include key prototype selection, cumulative routing quality allocation, and Gram regularization to enhance routing sparsity and diversity, revealing a negative correlation between shared coverage and residual demand. Experiments demonstrate that UniF-MoE outperforms both static and dynamic MoE approaches on DomainBed and GLUE benchmarks while significantly reducing activated computation, inference latency, and memory footprint.

0 citationsRead paper
Recent publications

Latest Papers

Share First, Route What Remains: A Unified Framework for Token-Adaptive MoE Computation

Aug 10, 2026

This work addresses the disconnect in existing Mixture-of-Experts (MoE) models between shared computation and dynamic routing, which overlooks the interdependence between reusable computation and residual expert requirements. The paper proposes UniF-MoE, a unified framework introducing a novel “shared-first, routed-later” mechanism: it first processes common features through a shared general-purpose module and then dynamically activates residual experts based on a shared-demand score and complementarity. Key innovations include key prototype selection, cumulative routing quality allocation, and Gram regularization to enhance routing sparsity and diversity, revealing a negative correlation between shared coverage and residual demand. Experiments demonstrate that UniF-MoE outperforms both static and dynamic MoE approaches on DomainBed and GLUE benchmarks while significantly reducing activated computation, inference latency, and memory footprint.

0 citationsRead paper

FreeShadow: Training-Free Shadow Removal via Illumination Transfer and Selective Content Preservation in Diffusion Models

Jul 29, 2026

This work addresses the limitations of existing shadow removal methods, which often suffer from poor generalization or artifacts due to insufficient training data diversity or time-consuming test-time optimization. The authors propose a training-free, optimization-free zero-shot approach leveraging a pre-trained diffusion model. By integrating an Illumination Transfer Attention (ITA) mechanism and a Local Texture-Preserving Relighting (LTPR) strategy, the method simultaneously achieves accurate illumination recovery and content fidelity without fine-tuning. It exploits the diffusion model’s self-attention maps and latent high-frequency features to effectively preserve local textures while transferring plausible lighting. The approach generates realistic, shadow-free images across diverse scenarios and significantly outperforms both current zero-shot and supervised methods.

0 citationsRead paper

P\textsuperscript{2}-DPO: Grounding Hallucination in Perceptual Processing via Calibration Direct Preference Optimization

Jun 02, 2026

This work addresses the perceptual bottlenecks commonly observed in large vision-language models, which often manifest as attentional bias and insufficient robustness under image degradation. To mitigate these issues, the authors propose P²-DPO, a novel training paradigm built upon the Direct Preference Optimization (DPO) framework. P²-DPO introduces, for the first time, a perception-oriented online self-generated preference pair mechanism and incorporates a calibration loss to achieve causal alignment between vision and language modalities. Notably, this approach operates without human feedback and, at comparable training cost, substantially enhances model performance in terms of attention region fidelity and robustness in degraded visual conditions, thereby improving both perceptual accuracy and visual reliability.

0 citationsRead paper