Institution profile

GAC R&D Center

Industry researchasia · cn
Official website
Research library8linked papers
Opportunities0open roles
Selected work

Representative Papers

SinD 2.0: A Multi-City UAV Dataset with Semantic Risk Annotations for SOTIF-Oriented Safety Validation at Signalized Intersections

Jul 18, 2026

This work addresses the safety validation challenges faced by autonomous driving systems at signalized intersections—particularly due to high traffic heterogeneity, sparsity of safety-critical events, lack of semantic risk annotations, and geographic homogeneity—by introducing the first large-scale, multi-city intersection dataset captured via drones across six intersections in four Chinese cities. The dataset includes 32,682 densely sampled edge-case scenarios and features a hierarchical semantic risk annotation framework encompassing traffic violations, visual occlusions, and narrow drivable areas. Integrated with SPaT (Signal Phase and Timing) data and high-definition maps, it supports both open- and closed-loop simulation testing. Experimental results reveal significant cross-domain distribution shifts within the dataset, and its semantically annotated risk subsets effectively expose algorithmic vulnerabilities, thereby establishing a high-quality benchmark for SOTIF (Safety of the Intended Functionality) validation.

0 citationsRead paper

Head Stabilization for Wheeled Bipedal Robots via Force-Estimation-Based Admittance Control

Nov 23, 2025

When traversing uneven terrain, wheeled bipedal robots exhibit vertical head oscillations in the world frame due to ground-induced disturbances, degrading onboard sensor accuracy and risking payload damage. To address this, we propose a model-based ground contact force estimation algorithm integrated with an admittance control strategy, enabling— for the first time—active head stabilization of wheeled bipedal robots in the world coordinate frame. Our approach leverages a 6-DOF dynamic model to estimate ground reaction forces online and dynamically regulate head orientation in real time. Simulation results demonstrate millisecond-level computational latency for force estimation; head displacement fluctuations are reduced by 82%. The system exhibits high robustness and superior dynamic response across sloped, stepped, and randomly irregular terrains, significantly enhancing terrain adaptability and perceptual reliability.

0 citationsRead paper

QA-VLM: Providing human-interpretable quality assessment for wire-feed laser additive manufacturing parts with Vision Language Models

Aug 20, 2025

In metal additive manufacturing, quality assessment remains heavily reliant on expert experience, while existing AI-based approaches lack interpretability. Method: This paper proposes the first explainable quality assessment framework integrating vision-language models (VLMs) with domain knowledge. It distills metallurgical expertise from academic literature into a VLM and employs attention mechanisms to enable semantic-level defect reasoning and natural-language explanation generation. Contribution/Results: Evaluated on 24 single-bead laser-wire deposition samples, the framework achieves significant improvements over generic VLMs in both assessment accuracy (+12.7% F1-score) and explanation consistency (+18.3% BLEU-4). By grounding visual reasoning in domain semantics and generating human-readable justifications, it enhances transparency and operational utility—directly addressing industry’s dual requirements for interpretability and practicality.

0 citationsRead paper

Decoupled Diffusion Sparks Adaptive Scene Generation

Apr 14, 2025

Existing traffic-scene generation methods face two key bottlenecks: full-sequence denoising compromises online responsiveness, while frame-wise prediction lacks explicit object-state guidance; moreover, open datasets predominantly cover routine behaviors, hindering realistic generation of high-risk corner cases. This paper proposes Nexus, a decoupled diffusion framework that introduces partial noise masking training and noise-aware scheduling—enabling disentangled modeling of sequential coherence and scenario challenge during layout generation. We further design fine-grained tokenized diffusion, independent noise-state modeling for dynamic agents, and closed-loop planning co-optimization. To support rigorous evaluation, we construct the first 540-hour high-risk corner-case simulation dataset. Experiments demonstrate a 40% reduction in displacement error and a 20% improvement in closed-loop planning performance, significantly enhancing realism and safety in complex interactive scenarios—including aggressive cut-ins, emergency braking, and collision avoidance.

0 citationsRead paper
Recent publications

Latest Papers

SinD 2.0: A Multi-City UAV Dataset with Semantic Risk Annotations for SOTIF-Oriented Safety Validation at Signalized Intersections

Jul 18, 2026

This work addresses the safety validation challenges faced by autonomous driving systems at signalized intersections—particularly due to high traffic heterogeneity, sparsity of safety-critical events, lack of semantic risk annotations, and geographic homogeneity—by introducing the first large-scale, multi-city intersection dataset captured via drones across six intersections in four Chinese cities. The dataset includes 32,682 densely sampled edge-case scenarios and features a hierarchical semantic risk annotation framework encompassing traffic violations, visual occlusions, and narrow drivable areas. Integrated with SPaT (Signal Phase and Timing) data and high-definition maps, it supports both open- and closed-loop simulation testing. Experimental results reveal significant cross-domain distribution shifts within the dataset, and its semantically annotated risk subsets effectively expose algorithmic vulnerabilities, thereby establishing a high-quality benchmark for SOTIF (Safety of the Intended Functionality) validation.

0 citationsRead paper

Head Stabilization for Wheeled Bipedal Robots via Force-Estimation-Based Admittance Control

Nov 23, 2025

When traversing uneven terrain, wheeled bipedal robots exhibit vertical head oscillations in the world frame due to ground-induced disturbances, degrading onboard sensor accuracy and risking payload damage. To address this, we propose a model-based ground contact force estimation algorithm integrated with an admittance control strategy, enabling— for the first time—active head stabilization of wheeled bipedal robots in the world coordinate frame. Our approach leverages a 6-DOF dynamic model to estimate ground reaction forces online and dynamically regulate head orientation in real time. Simulation results demonstrate millisecond-level computational latency for force estimation; head displacement fluctuations are reduced by 82%. The system exhibits high robustness and superior dynamic response across sloped, stepped, and randomly irregular terrains, significantly enhancing terrain adaptability and perceptual reliability.

0 citationsRead paper

QA-VLM: Providing human-interpretable quality assessment for wire-feed laser additive manufacturing parts with Vision Language Models

Aug 20, 2025

In metal additive manufacturing, quality assessment remains heavily reliant on expert experience, while existing AI-based approaches lack interpretability. Method: This paper proposes the first explainable quality assessment framework integrating vision-language models (VLMs) with domain knowledge. It distills metallurgical expertise from academic literature into a VLM and employs attention mechanisms to enable semantic-level defect reasoning and natural-language explanation generation. Contribution/Results: Evaluated on 24 single-bead laser-wire deposition samples, the framework achieves significant improvements over generic VLMs in both assessment accuracy (+12.7% F1-score) and explanation consistency (+18.3% BLEU-4). By grounding visual reasoning in domain semantics and generating human-readable justifications, it enhances transparency and operational utility—directly addressing industry’s dual requirements for interpretability and practicality.

0 citationsRead paper

Decoupled Diffusion Sparks Adaptive Scene Generation

Apr 14, 2025

Existing traffic-scene generation methods face two key bottlenecks: full-sequence denoising compromises online responsiveness, while frame-wise prediction lacks explicit object-state guidance; moreover, open datasets predominantly cover routine behaviors, hindering realistic generation of high-risk corner cases. This paper proposes Nexus, a decoupled diffusion framework that introduces partial noise masking training and noise-aware scheduling—enabling disentangled modeling of sequential coherence and scenario challenge during layout generation. We further design fine-grained tokenized diffusion, independent noise-state modeling for dynamic agents, and closed-loop planning co-optimization. To support rigorous evaluation, we construct the first 540-hour high-risk corner-case simulation dataset. Experiments demonstrate a 40% reduction in displacement error and a 20% improvement in closed-loop planning performance, significantly enhancing realism and safety in complex interactive scenarios—including aggressive cut-ins, emergency braking, and collision avoidance.

0 citationsRead paper