Institution profile

Shanghai University

Academic institutionasia · cn
Official website
Research library358linked papers
Opportunities0open roles
Selected work

Representative Papers

Large Model Empowered Metaverse: State-of-the-Art, Challenges and Opportunities

Jan 18, 2025arXiv.org

Metaverse applications face critical bottlenecks including high real-time rendering latency, poor adaptability to dynamic scenes, and limited scalability. To address these challenges, this paper proposes a large language model (LLM)-empowered cloud-edge-device collaborative generative AI rendering framework. It introduces two key innovations: (1) a mobility-aware pre-rendering mechanism that anticipates user movement for proactive resource allocation, and (2) a diffusion model–driven adaptive rendering strategy that dynamically optimizes visual fidelity and computational load based on scene complexity and device capabilities. The framework tightly integrates LLMs, video foundation models (e.g., Sora), and hierarchical distributed computing across cloud, edge, and end devices. Experimental evaluation demonstrates a 37% reduction in end-to-end rendering latency and significantly enhanced real-time immersion under high-concurrency, highly dynamic conditions. This work establishes a scalable, generative-AI-native technical pathway for next-generation metaverse systems.

3 citationsRead paper

Tumor Detection, Segmentation and Classification Challenge on Automated 3D Breast Ultrasound: The TDSC-ABUS Challenge

Jan 26, 2025

Automatic breast ultrasound (ABUS) tumor detection, segmentation, and classification are challenged by morphological heterogeneity, low signal-to-noise ratio, and scarcity of annotated 3D data. Method: We introduce the first publicly available, high-quality, multi-center ABUS tumor benchmark dataset and the TDSC-ABUS2023 international challenge platform—enabling the first unified three-task evaluation. Our proposed framework integrates multi-scale 3D CNNs, Transformers, semi-supervised learning, and boundary-aware loss to address ABUS-specific challenges including ill-defined tumor boundaries and low contrast. Contribution/Results: Our method achieves state-of-the-art performance: 82.3% mAP@0.5 for detection, 79.6% Dice for segmentation, and 91.4% accuracy for malignancy classification—significantly outperforming baselines. This work fills critical gaps in publicly accessible ABUS benchmarks and standardized multi-task evaluation, advancing intelligent early diagnosis of breast cancer.

2 citationsRead paper

MTAVG-Bench: A Comprehensive Benchmark for Evaluating Multi-Talker Dialogue-Centric Audio-Video Generation

Jan 31, 2026

Existing evaluation benchmarks struggle to effectively assess critical issues in multi-speaker conversational video generation, such as identity drift, unnatural turn-taking, and audio-visual asynchrony. This work proposes the first fine-grained audiovisual generation evaluation framework tailored to this scenario, introducing a comprehensive benchmark comprising 1.8K videos and 2.4K structured question-answer pairs, constructed via a semi-automatic pipeline. The framework evaluates models across four dimensions: audiovisual fidelity, temporal consistency, social interaction coherence, and cinematic expressiveness, enabling precise failure analysis and targeted model refinement. Experiments on twelve leading open- and closed-source models reveal that Gemini 3 Pro achieves the best overall performance, while certain open-source models excel in signal fidelity and temporal consistency.

1 citationsRead paper

Efficient UAV trajectory prediction: A multi-modal deep diffusion framework

Jan 26, 2026

This work addresses the challenge of insufficient accuracy in predicting trajectories of unauthorized drones in low-altitude airspace, which stems from the limited information provided by single-sensor systems. To overcome this limitation, the authors propose a multimodal deep diffusion framework that fuses point clouds from LiDAR and millimeter-wave radar. The approach employs structurally aligned dual-branch encoders to extract modality-specific features and introduces a bidirectional cross-attention mechanism to achieve semantic alignment and complementary fusion of geometric structures and dynamic reflectivity characteristics. A tailored loss function and post-processing strategy are further integrated to enhance prediction performance. Evaluated on the MMAUD dataset, the proposed method achieves a 40% improvement in trajectory prediction accuracy over baseline models, demonstrating the effectiveness and practicality of the multimodal fusion strategy.

1 citationsRead paper

TacUMI: A Multi-Modal Universal Manipulation Interface for Contact-Rich Tasks

Jan 21, 2026

This work addresses the challenge of accurately identifying semantic event boundaries in long-horizon manipulation tasks rich in physical contact, where reliance solely on visual and proprioceptive cues proves insufficient for effective task segmentation. To overcome this limitation, the authors propose TacUMI—a compact, multimodal data acquisition system that integrates ViTac visuo-tactile sensing with force-torque and pose perception—and, for the first time, embed it within a general-purpose manipulation interface to enable highly synchronized multimodal recording. Building upon this hardware foundation, they further introduce a temporal modeling–based multimodal fusion framework to automatically extract event boundaries from human demonstrations. Evaluated on a cable assembly task, the method achieves over 90% segmentation accuracy, significantly outperforming unimodal baselines and demonstrating the critical role of multimodal perception in enhancing task decomposition performance.

1 citationsRead paper
Recent publications

Latest Papers

Newton Deep Unfolding for Compressed Sensing

Sep 13, 2026

为解决压缩感知中测量数据有限及现有方法优化状态利用不足的问题,提出了一种基于二阶优化的Newton深度展开网络(NDU-Net),通过引入Newton更新模块和多尺度先验模块提升重建性能。

0 citationsRead paper