Institution profile

Great Wall Motor Co., Ltd.

Industry researchasia · cn
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

DriveCache: Action-Aware Caching for Driving World Model Inference

Aug 17, 2026

This study addresses the inference bottleneck in driving video generation models caused by redundant computations during diffusion. We propose DriveCache, a training-free, action-aware cache controller that pioneers the integration of driving signals into cache scheduling. By optimizing denoising steps through dynamic programming and employing causal drift detection for adaptive feature reuse and correction, DriveCache significantly enhances both inference efficiency and generation fidelity within responsive computational budgets. Experimental evaluations across three generator configurations demonstrate that DriveCache achieves superior efficiency-quality trade-offs compared to existing caching methods. Consequently, this work establishes a novel paradigm for the efficient deployment of world models in autonomous driving applications, effectively balancing computational constraints with high-fidelity video synthesis without requiring additional model retraining.

0 citationsRead paper

Visual Geometry Foundation-Aware Gaussians for Single-Frame Surround-View Driving Reconstruction

Aug 11, 2026

This work addresses the geometric instability and rendering artifacts commonly encountered in single-frame surround-view driving scene reconstruction, which stem from sparse overlap among camera views. To mitigate these issues, the authors propose the VGGD framework, which introduces visual geometric foundation model priors into this task for the first time. Specifically, a Visual Geometric Foundation Tokenizer (VGGT) generates multi-view geometric prior tokens to enhance front-end geometry modeling. A dual-path neck network is designed to disentangle geometric and appearance representations, complemented by a scale-warmup strategy and a hybrid pixel-voxel Gaussian decoder. Evaluated on the single-frame nuScenes benchmark, the proposed method significantly improves geometric consistency and rendering quality, outperforming existing approaches.

0 citationsRead paper

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects

Aug 08, 2026

This work addresses the challenges of data scarcity and misalignment between semantic representations and fixed acoustic supervision in end-to-end spoken dialogue modeling for low-resource Chinese dialects. To overcome these issues, the authors construct a scalable data pipeline for dialectal spoken dialogue synthesis and propose a two-stage post-training strategy incorporating a self-aligned speech supervision mechanism that dynamically aligns acoustic targets with the model’s evolving semantic representations. This approach achieves the first successful end-to-end spoken dialogue modeling for low-resource Chinese dialects, significantly outperforming existing baselines across multiple dialects. Substantial improvements are observed in dialect consistency, response quality, and speech intelligibility. The complete framework is open-sourced to facilitate future research in this underexplored domain.

0 citationsRead paper

Depth-Copy-Paste: Multimodal and Depth-Aware Compositing for Robust Face Detection

Dec 12, 2025

Traditional copy-paste data augmentation for face detection often yields distorted synthetic images due to inaccurate foreground segmentation, geometric misalignment, and semantically inconsistent backgrounds. To address these issues, we propose a depth-guided multimodal synthesis framework. Our method introduces a novel depth-guided sliding-window pasting mechanism to ensure physical plausibility; integrates BLIP and CLIP for cross-modal semantic-visual joint retrieval; and leverages SAM3 and Depth-Anything to extract occlusion-free, visible human regions—preserving facial texture fidelity while optimizing depth map alignment. This approach significantly enhances the robustness of face detectors under challenging conditions, including heavy occlusion, low illumination, and complex backgrounds. On WIDER FACE, it achieves absolute mAP gains of 3.2–5.7% over state-of-the-art depth-free augmentation methods.

0 citationsRead paper
Recent publications

Latest Papers

DriveCache: Action-Aware Caching for Driving World Model Inference

Aug 17, 2026

This study addresses the inference bottleneck in driving video generation models caused by redundant computations during diffusion. We propose DriveCache, a training-free, action-aware cache controller that pioneers the integration of driving signals into cache scheduling. By optimizing denoising steps through dynamic programming and employing causal drift detection for adaptive feature reuse and correction, DriveCache significantly enhances both inference efficiency and generation fidelity within responsive computational budgets. Experimental evaluations across three generator configurations demonstrate that DriveCache achieves superior efficiency-quality trade-offs compared to existing caching methods. Consequently, this work establishes a novel paradigm for the efficient deployment of world models in autonomous driving applications, effectively balancing computational constraints with high-fidelity video synthesis without requiring additional model retraining.

0 citationsRead paper

Visual Geometry Foundation-Aware Gaussians for Single-Frame Surround-View Driving Reconstruction

Aug 11, 2026

This work addresses the geometric instability and rendering artifacts commonly encountered in single-frame surround-view driving scene reconstruction, which stem from sparse overlap among camera views. To mitigate these issues, the authors propose the VGGD framework, which introduces visual geometric foundation model priors into this task for the first time. Specifically, a Visual Geometric Foundation Tokenizer (VGGT) generates multi-view geometric prior tokens to enhance front-end geometry modeling. A dual-path neck network is designed to disentangle geometric and appearance representations, complemented by a scale-warmup strategy and a hybrid pixel-voxel Gaussian decoder. Evaluated on the single-frame nuScenes benchmark, the proposed method significantly improves geometric consistency and rendering quality, outperforming existing approaches.

0 citationsRead paper

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects

Aug 08, 2026

This work addresses the challenges of data scarcity and misalignment between semantic representations and fixed acoustic supervision in end-to-end spoken dialogue modeling for low-resource Chinese dialects. To overcome these issues, the authors construct a scalable data pipeline for dialectal spoken dialogue synthesis and propose a two-stage post-training strategy incorporating a self-aligned speech supervision mechanism that dynamically aligns acoustic targets with the model’s evolving semantic representations. This approach achieves the first successful end-to-end spoken dialogue modeling for low-resource Chinese dialects, significantly outperforming existing baselines across multiple dialects. Substantial improvements are observed in dialect consistency, response quality, and speech intelligibility. The complete framework is open-sourced to facilitate future research in this underexplored domain.

0 citationsRead paper

Depth-Copy-Paste: Multimodal and Depth-Aware Compositing for Robust Face Detection

Dec 12, 2025

Traditional copy-paste data augmentation for face detection often yields distorted synthetic images due to inaccurate foreground segmentation, geometric misalignment, and semantically inconsistent backgrounds. To address these issues, we propose a depth-guided multimodal synthesis framework. Our method introduces a novel depth-guided sliding-window pasting mechanism to ensure physical plausibility; integrates BLIP and CLIP for cross-modal semantic-visual joint retrieval; and leverages SAM3 and Depth-Anything to extract occlusion-free, visible human regions—preserving facial texture fidelity while optimizing depth map alignment. This approach significantly enhances the robustness of face detectors under challenging conditions, including heavy occlusion, low illumination, and complex backgrounds. On WIDER FACE, it achieves absolute mAP gains of 3.2–5.7% over state-of-the-art depth-free augmentation methods.

0 citationsRead paper