Institution profile

Fudan University

Academic institutionasia · cn
Official website
Research library3,420linked papers
Opportunities0open roles
Selected work

Representative Papers

Efficient4D: Fast Dynamic 3D Object Generation from a Single-view Video

Jan 16, 2024

To address the challenges of missing 4D annotations and low end-to-end optimization efficiency in monocular video-based dynamic 3D reconstruction, this work proposes a two-stage decoupled paradigm: first generating multi-view temporally consistent images via a diffusion model, then driving 4D Gaussian Splatting for explicit reconstruction. We introduce an inconsistency-aware confidence-weighted loss and a lightweight Score Distillation Sampling (SDS) loss, significantly improving robustness under sparse-view conditions. Compared to Consistent4D, our method accelerates training tenfold (10 minutes vs. 120 minutes), enables real-time continuous trajectory rendering, and achieves state-of-the-art novel-view synthesis quality. To the best of our knowledge, this is the first work to organically integrate generative modeling with explicit 4D reconstruction, establishing a new paradigm for efficient, high-fidelity dynamic scene reconstruction.

42 citations5 influentialRead paper

MeViS: A Multi-Modal Dataset for Referring Motion Expression Video Segmentation

Aug 19, 2025IEEE Transactions on Pattern Analysis and Machine Intelligence

Existing video segmentation datasets emphasize static attribute descriptions, neglecting the critical role of motion in video understanding. To address this, we introduce MeViS—the first multimodal video segmentation dataset explicitly guided by motion expression—comprising 33K human-annotated text and audio motion descriptions across 2,006 complex scenes and 8,171 objects, supporting four tasks: Referring Video Object Segmentation (RVOS), Audio-Visual Object Segmentation (AVOS), Referring Multi-Object Tracking (RMOT), and Referring Motion Expression Grounding (RMEG). MeViS pioneers motion semantics as the core referential cue, breaking the static-dominant paradigm. We further propose LMPM++, a model integrating multimodal aligned annotation, motion-aware modeling, and joint audio-visual-linguistic representation, achieving new state-of-the-art performance on RVOS, AVOS, and RMOT. Comprehensive evaluation of 15 mainstream methods reveals systematic motion reasoning bottlenecks; leveraging MeViS significantly improves segmentation and tracking accuracy, advancing motion-centric video understanding.

22 citationsRead paper

SlowFocus: Enhancing Fine-grained Temporal Understanding in Video LLM

Feb 03, 2026Neural Information Processing Systems

Existing video large language models struggle to simultaneously preserve frame-level semantic details and capture video-level temporal structure, limiting their fine-grained understanding capabilities. To address this challenge, this work proposes SlowFocus, a mechanism that identifies question-relevant temporal segments and applies dense sampling, integrated with a multi-band hybrid attention module to effectively fuse local high-frequency visual details with global low-frequency contextual information. This approach significantly enhances the effective sampling rate without compromising the quality of frame-level visual tokens. Additionally, we introduce a training strategy tailored for fine-grained temporal reasoning and construct a new benchmark, FineAction-CGR. Extensive experiments demonstrate consistent and substantial performance gains across multiple established video understanding benchmarks as well as FineAction-CGR, confirming the superiority of our method in fine-grained temporal understanding tasks.

17 citationsRead paper

A Queueing Theoretic Perspective on Low-Latency LLM Inference with Variable Token Length

Jul 07, 2024International Symposium on Modeling and Optimization in Mobile, Ad-Hoc and Wireless Networks

Variable-length outputs in LLM interactive serving induce significant inference queuing latency due to output-token–dependent service times. Method: We propose a unified theoretical framework integrating M/G/1 and batch-service queueing models, the first to treat output token count as a stochastic service time. We jointly optimize the max-token limit and batch scheduling policies—fixed, dynamic, and elastic—to characterize their distinct latency behaviors under output-length uncertainty. Contribution/Results: Our analysis reveals the dominant impact of long-tail requests on mean queuing delay. Event-driven simulations validate model accuracy (<5% error): setting max-token = 256 reduces mean queuing delay by 38%; under load fluctuations, elastic batching cuts delay by 22% versus fixed batching. The core contribution is establishing a quantitative relationship between output-length variability and system latency, enabling principled co-optimization of inference parameters.

12 citationsRead paper

ForgerySleuth: Empowering Multimodal Large Language Models for Image Manipulation Detection

Nov 29, 2024arXiv.org

Existing methods have not explored the potential of multimodal large language models (M-LLMs) for image forgery detection; direct application often induces hallucination and overthinking, resulting in inaccurate reasoning and coarse-grained localization. Method: We propose the first Chain-of-Clues prompting paradigm tailored for forgery detection, accompanied by theForgeryAnalysis dataset; design a joint clue-fusion and segmentation-generation architecture enabling text-guided pixel-level localization; and introduce a data synthesis and augmentation engine to support large-scale pretraining. Contribution/Results: Our approach significantly improves generalizability, robustness, and interpretability. It outperforms state-of-the-art methods across multiple benchmarks and achieves, for the first time, end-to-end, interpretable, fine-grained M-LLM–driven forgery analysis.

10 citations1 influentialRead paper
Recent publications

Latest Papers