Institution profile

Huawei

Industry researchasia · cn
Official website
Research library2,069linked papers
Opportunities0open roles
Selected work

Representative Papers

Depth and Image Fusion for Road Obstacle Detection Using Stereo Camera

Apr 11, 20262

Real-time detection of small obstacles (e.g., hubcaps, cardboard boxes) under complex illumination and unstructured road surfaces remains challenging due to low contrast, ambiguous textures, and lack of prior knowledge. Method: This paper proposes a training-free depth-RGB collaborative detection framework. It fuses stereo-derived depth maps with RGB imagery via a multimodal superpixel fusion mechanism, jointly enhancing SLIC-guided stereo matching and fine-grained texture discrimination for small objects. The approach operates without scene priors or annotated data. Contribution/Results: The method achieves robust detection and stable tracking of static and low-speed small obstacles of arbitrary size, shape, and appearance time. Evaluated in underground parking lots, it significantly improves recall under low-contrast and dynamically varying artificial lighting conditions. By eliminating reliance on labeled datasets or domain-specific assumptions, it offers a cost-effective, highly adaptive perception solution for autonomous driving systems.

2,026 citations2,026 influentialRead paper

SlowFocus: Enhancing Fine-grained Temporal Understanding in Video LLM

Feb 03, 2026Neural Information Processing Systems

Existing video large language models struggle to simultaneously preserve frame-level semantic details and capture video-level temporal structure, limiting their fine-grained understanding capabilities. To address this challenge, this work proposes SlowFocus, a mechanism that identifies question-relevant temporal segments and applies dense sampling, integrated with a multi-band hybrid attention module to effectively fuse local high-frequency visual details with global low-frequency contextual information. This approach significantly enhances the effective sampling rate without compromising the quality of frame-level visual tokens. Additionally, we introduce a training strategy tailored for fine-grained temporal reasoning and construct a new benchmark, FineAction-CGR. Extensive experiments demonstrate consistent and substantial performance gains across multiple established video understanding benchmarks as well as FineAction-CGR, confirming the superiority of our method in fine-grained temporal understanding tasks.

17 citationsRead paper

Your Absorbing Discrete Diffusion Secretly Models the Conditional Distributions of Clean Data

Jun 06, 2024arXiv.org

This work addresses the inefficiency of estimating marginal probability ratios—termed “concrete scores”—in absorption-based discrete diffusion models. We establish, for the first time, an analytical equivalence between concrete scores and clean-data conditional probabilities: concrete scores decompose into the product of the conditional probability and a closed-form time-dependent factor. Leveraging this insight, we propose RADD (Time-Invariant Reparameterized Absorption Diffusion), a reparameterization that eliminates explicit dependence on timestep indices and enables NFE (number-of-function-evaluations) caching for accelerated sampling. Theoretically, our framework unifies performance bounds for absorption diffusion and arbitrary-order autoregressive models. Empirically, RADD achieves state-of-the-art perplexity among diffusion-based language models at the GPT-2 scale across five zero-shot language modeling benchmarks. Code is publicly available.

9 citations2 influentialRead paper

Gradually Excavating External Knowledge for Implicit Complex Question Answering

Mar 09, 2026Conference on Empirical Methods in Natural Language Processing

Large language models often struggle with open-domain implicit complex reasoning due to limited knowledge coverage, poor temporal relevance, and insufficient reasoning caused by one-shot generation. To address these limitations, this work proposes a progressive external knowledge mining framework that dynamically selects between knowledge retrieval and logical reasoning at each iterative step through an adaptive action selection mechanism, thereby effectively integrating external knowledge with multi-step reasoning. Evaluated on the StrategyQA benchmark, the proposed method achieves an accuracy of 78.17% with less than 6% of the parameter count of competing models, establishing a new state-of-the-art performance among models of the 10B scale.

7 citationsRead paper

SWE-Lego: Pushing the Limits of Supervised Fine-tuning for Software Issue Resolving

Jan 04, 2026arXiv.org

This work explores an efficient, lightweight paradigm for addressing software engineering tasks using only supervised fine-tuning (SFT), without relying on reinforcement learning or complex alignment techniques. To this end, we construct a high-quality hybrid dataset combining real-world and synthetically generated samples, and introduce several novel components: an error-masking mechanism, a software engineering–oriented curriculum learning strategy based on task difficulty, and a test-time scaling (TTS) approach integrated with trajectory validation. Our method achieves state-of-the-art performance among open-source models on SWE-bench Verified: SWE-Lego-Qwen3-8B and SWE-Lego-Qwen3-32B attain pass rates of 42.2% and 52.6%, respectively, which further improve to 49.6% and 58.8% under TTS@16.

5 citations2 influentialRead paper
Recent publications

Latest Papers