Institution profile

Xi'an Jiao Tong University

Academic institutionasia · cn
Official website
Research library1,576linked papers
Opportunities0open roles
Selected work

Representative Papers

Selective Transfer Learning of Cross-Modality Distillation for Monocular 3D Object Detection

Oct 01, 2024IEEE transactions on circuits and systems for video technology (Print)

Monocular 3D object detection is inherently ill-posed due to the absence of precise depth information, and existing cross-modal knowledge distillation approaches often suffer from negative transfer caused by the modality gap between images and LiDAR. To address this issue, this work proposes MonoSTL, which presents the first systematic analysis of negative transfer in cross-modal distillation and introduces two novel components: Depth-Aware Selective Feature Distillation (DASFD) and Depth-Aware Selective Relation Distillation (DASRD). These modules leverage depth uncertainty to guide positive knowledge transfer and effectively integrate LiDAR-derived depth cues through structural alignment and selective distillation mechanisms. Extensive experiments demonstrate that MonoSTL significantly boosts the performance of various baseline models on both KITTI and NuScenes benchmarks, achieving state-of-the-art results and confirming its effectiveness and generalizability.

11 citationsRead paper

Instance-Conditioned Adaptation for Large-scale Generalization of Neural Combinatorial Optimization

May 03, 2024arXiv.org

Existing neural combinatorial optimization (NCO) methods exhibit poor generalization to large-scale routing problems—such as the Traveling Salesman Problem (TSP) and Capacitated Vehicle Routing Problem (CVRP)—limiting their applicability in real-world intelligent transportation systems. To address this, we propose Instance-Conditional Adaptive Mechanism (ICAM), a construction-based graph neural network model that achieves cross-scale adaptability via lightweight adapters conditioned on instance-specific embeddings. We further introduce a novel three-stage unsupervised reinforcement learning paradigm, enabling end-to-end training on instances ranging from 100 to 1,000 nodes without access to optimal solution labels. Experiments demonstrate that ICAM achieves state-of-the-art performance among construction-based NCO approaches on TSP and CVRP benchmarks, scales robustly up to 1,000 nodes, and delivers highly efficient inference—significantly outperforming existing methods.

6 citations1 influentialRead paper

ViLa-MIL: Dual-scale Vision-Language Multiple Instance Learning for Whole Slide Image Classification

Jun 16, 2024Computer Vision and Pattern Recognition

In whole-slide image (WSI) classification for digital pathology, existing multiple instance learning (MIL) methods suffer from heavy reliance on abundant bag-level annotations and poor generalizability, while vision-language models (VLMs) are hindered by pathology-agnostic text prompts and prohibitively high pretraining costs, yielding limited performance gains. To address these limitations, we propose ViLa-MIL—a dual-scale vision-language MIL framework. It introduces the first pathology-informed, dual-scale descriptive text prompting mechanism; designs a prototype-guided patch decoder and a context-guided text decoder to enable cross-modal, multi-granularity feature co-modeling; and integrates a frozen large language model, prototype clustering, and vision-language alignment. Evaluated on three multi-cancer, multi-center datasets, ViLa-MIL significantly outperforms state-of-the-art methods, demonstrating low annotation dependency, strong cross-center generalizability, and high robustness.

4 citations1 influentialRead paper

Matching Distance and Geometric Distribution Aided Learning Multiview Point Cloud Registration

Nov 01, 2024IEEE Robotics and Automation Letters

To address the unreliable pose graph construction and motion synchronization challenges in multi-view point cloud registration, this paper proposes an end-to-end absolute pose estimation paradigm. First, matching distance is introduced as a principled reliability metric for pose graph construction, replacing handcrafted loss functions with direct global pose regression. Second, the method jointly optimizes feature interaction and structural awareness by integrating local geometric distribution modeling with adaptive attention mechanisms. Fully data-driven, it eliminates iterative optimization and post-processing. Evaluated on diverse indoor and outdoor datasets, the approach achieves a 12.7% improvement in pose graph construction accuracy and reduces overall registration error by 21.3%, demonstrating significantly enhanced robustness and cross-scene generalization capability.

3 citations1 influentialRead paper

Large Model Empowered Metaverse: State-of-the-Art, Challenges and Opportunities

Jan 18, 2025arXiv.org

Metaverse applications face critical bottlenecks including high real-time rendering latency, poor adaptability to dynamic scenes, and limited scalability. To address these challenges, this paper proposes a large language model (LLM)-empowered cloud-edge-device collaborative generative AI rendering framework. It introduces two key innovations: (1) a mobility-aware pre-rendering mechanism that anticipates user movement for proactive resource allocation, and (2) a diffusion model–driven adaptive rendering strategy that dynamically optimizes visual fidelity and computational load based on scene complexity and device capabilities. The framework tightly integrates LLMs, video foundation models (e.g., Sora), and hierarchical distributed computing across cloud, edge, and end devices. Experimental evaluation demonstrates a 37% reduction in end-to-end rendering latency and significantly enhanced real-time immersion under high-concurrency, highly dynamic conditions. This work establishes a scalable, generative-AI-native technical pathway for next-generation metaverse systems.

3 citationsRead paper
Recent publications

Latest Papers