Institution profile

OPPO Research Institute

Industry researchasia · cn
Official website
Research library186linked papers
Opportunities0open roles
Selected work

Representative Papers

ANUBIS: Skeleton Action Recognition Dataset, Review, and Benchmark

May 04, 2022arXiv.org

Existing 3D skeleton-based action recognition research suffers from fragmented representation taxonomies and evaluation protocols misaligned with real-world scenarios; moreover, mainstream datasets lack critical dimensions—including rear-view perspectives, multi-person interactions, fine-grained or violent actions, and pandemic-era behaviors. To address these gaps, we propose a four-dimensional taxonomy (dataset design, spatial modeling, temporal modeling, and signal enhancement) and introduce ANUBIS: the first large-scale, multi-view 3D skeleton dataset explicitly designed for realistic challenges. ANUBIS features rear-view captures, 101 action classes (including 21 pandemic-related behaviors), and standardized recordings from 128 participants using Azure Kinect’s multi-sensor fusion. We further establish a unified benchmark framework, enabling reproducible evaluation of 12 state-of-the-art models. Our analysis identifies temporal modeling capacity and signal robustness as the primary bottlenecks limiting current performance.

4 citationsRead paper

OpenRR-5k: A Large-Scale Benchmark for Reflection Removal in the Wild

Jun 05, 2025

Current single-image reflection removal (SIRR) research is hindered by the lack of large-scale, high-quality, real-world benchmark datasets. To address this, we introduce the first large-scale, in-the-wild SIRR benchmark—comprising 5,300 pixel-accurately aligned reflection/no-reflection image pairs—spanning diverse illumination conditions, object materials, and reflection patterns; it further includes 100 real-world, ground-truth-free images for generalization evaluation. A rigorously controlled acquisition pipeline ensures data fidelity. We propose an end-to-end U-Net-based removal model and comprehensively evaluate performance using five metrics: PSNR, SSIM, LPIPS, DISTS, and NIQE. Experiments confirm the dataset’s validity and establish a robust baseline. All data and code are publicly released to advance standardization and practical deployment of SIRR.

2 citationsRead paper

Degradation-Aware Image Enhancement via Vision-Language Classification

Jun 05, 2025

Real-world image degradations are diverse and challenging to identify automatically. Method: This paper pioneers the integration of vision-language models (VLMs) into degradation classification, proposing a fine-grained degradation-aware and modular restoration framework. It categorizes degradations into four types—super-resolution-related, reflection, motion blur, and no degradation—and employs a VLM to accurately classify each input image, subsequently activating a dedicated reconstruction network for on-demand enhancement. Domain adaptation is incorporated to improve cross-scenario generalization. Results: On multiple benchmark datasets, the method achieves 92.7% degradation classification accuracy and significantly outperforms unified enhancement approaches in PSNR and SSIM. It also enhances perceptual quality of restored images and improves performance on downstream tasks, thereby overcoming inherent limitations of end-to-end models in interpretability and generalizability.

2 citationsRead paper

Beyond External Guidance: Unleashing the Semantic Richness Inside Diffusion Transformers for Improved Training

Jan 12, 2026

Diffusion Transformers (DiTs) suffer from slow training convergence, and existing acceleration methods rely on external pretrained models, limiting their flexibility and generalization. This work proposes Self-Transcendence, a novel approach that dispenses with external semantic guidance and achieves fully self-supervised training by leveraging only internal model features. Specifically, it aligns shallow-layer DiT features with VAE latent representations and enhances semantic expressiveness of intermediate features through classifier-free guidance (CFG), enabling significant training acceleration using solely internal supervision signals. Experimental results demonstrate that, without any external pretrained models, the proposed method outperforms external-guidance approaches such as REPA in both training speed and generation quality.

1 citationsRead paper
Recent publications

Latest Papers