WAVE-Go: World-Model Navigation with Adaptive Execution for Wheel-Legged Robots
为解决轮腿机器人在动态障碍物或切换运动模式时导航问题,提出WAVE-Go框架,通过自适应执行和中断命令来提高导航成功率并减少碰撞。
为解决轮腿机器人在动态障碍物或切换运动模式时导航问题,提出WAVE-Go框架,通过自适应执行和中断命令来提高导航成功率并减少碰撞。
This work addresses the safety validation challenges faced by autonomous driving systems at signalized intersections—particularly due to high traffic heterogeneity, sparsity of safety-critical events, lack of semantic risk annotations, and geographic homogeneity—by introducing the first large-scale, multi-city intersection dataset captured via drones across six intersections in four Chinese cities. The dataset includes 32,682 densely sampled edge-case scenarios and features a hierarchical semantic risk annotation framework encompassing traffic violations, visual occlusions, and narrow drivable areas. Integrated with SPaT (Signal Phase and Timing) data and high-definition maps, it supports both open- and closed-loop simulation testing. Experimental results reveal significant cross-domain distribution shifts within the dataset, and its semantically annotated risk subsets effectively expose algorithmic vulnerabilities, thereby establishing a high-quality benchmark for SOTIF (Safety of the Intended Functionality) validation.
When traversing uneven terrain, wheeled bipedal robots exhibit vertical head oscillations in the world frame due to ground-induced disturbances, degrading onboard sensor accuracy and risking payload damage. To address this, we propose a model-based ground contact force estimation algorithm integrated with an admittance control strategy, enabling— for the first time—active head stabilization of wheeled bipedal robots in the world coordinate frame. Our approach leverages a 6-DOF dynamic model to estimate ground reaction forces online and dynamically regulate head orientation in real time. Simulation results demonstrate millisecond-level computational latency for force estimation; head displacement fluctuations are reduced by 82%. The system exhibits high robustness and superior dynamic response across sloped, stepped, and randomly irregular terrains, significantly enhancing terrain adaptability and perceptual reliability.
In metal additive manufacturing, quality assessment remains heavily reliant on expert experience, while existing AI-based approaches lack interpretability. Method: This paper proposes the first explainable quality assessment framework integrating vision-language models (VLMs) with domain knowledge. It distills metallurgical expertise from academic literature into a VLM and employs attention mechanisms to enable semantic-level defect reasoning and natural-language explanation generation. Contribution/Results: Evaluated on 24 single-bead laser-wire deposition samples, the framework achieves significant improvements over generic VLMs in both assessment accuracy (+12.7% F1-score) and explanation consistency (+18.3% BLEU-4). By grounding visual reasoning in domain semantics and generating human-readable justifications, it enhances transparency and operational utility—directly addressing industry’s dual requirements for interpretability and practicality.
Existing traffic-scene generation methods face two key bottlenecks: full-sequence denoising compromises online responsiveness, while frame-wise prediction lacks explicit object-state guidance; moreover, open datasets predominantly cover routine behaviors, hindering realistic generation of high-risk corner cases. This paper proposes Nexus, a decoupled diffusion framework that introduces partial noise masking training and noise-aware scheduling—enabling disentangled modeling of sequential coherence and scenario challenge during layout generation. We further design fine-grained tokenized diffusion, independent noise-state modeling for dynamic agents, and closed-loop planning co-optimization. To support rigorous evaluation, we construct the first 540-hour high-risk corner-case simulation dataset. Experiments demonstrate a 40% reduction in displacement error and a 20% improvement in closed-loop planning performance, significantly enhancing realism and safety in complex interactive scenarios—including aggressive cut-ins, emergency braking, and collision avoidance.
为解决轮腿机器人在动态障碍物或切换运动模式时导航问题,提出WAVE-Go框架,通过自适应执行和中断命令来提高导航成功率并减少碰撞。
This work addresses the safety validation challenges faced by autonomous driving systems at signalized intersections—particularly due to high traffic heterogeneity, sparsity of safety-critical events, lack of semantic risk annotations, and geographic homogeneity—by introducing the first large-scale, multi-city intersection dataset captured via drones across six intersections in four Chinese cities. The dataset includes 32,682 densely sampled edge-case scenarios and features a hierarchical semantic risk annotation framework encompassing traffic violations, visual occlusions, and narrow drivable areas. Integrated with SPaT (Signal Phase and Timing) data and high-definition maps, it supports both open- and closed-loop simulation testing. Experimental results reveal significant cross-domain distribution shifts within the dataset, and its semantically annotated risk subsets effectively expose algorithmic vulnerabilities, thereby establishing a high-quality benchmark for SOTIF (Safety of the Intended Functionality) validation.
When traversing uneven terrain, wheeled bipedal robots exhibit vertical head oscillations in the world frame due to ground-induced disturbances, degrading onboard sensor accuracy and risking payload damage. To address this, we propose a model-based ground contact force estimation algorithm integrated with an admittance control strategy, enabling— for the first time—active head stabilization of wheeled bipedal robots in the world coordinate frame. Our approach leverages a 6-DOF dynamic model to estimate ground reaction forces online and dynamically regulate head orientation in real time. Simulation results demonstrate millisecond-level computational latency for force estimation; head displacement fluctuations are reduced by 82%. The system exhibits high robustness and superior dynamic response across sloped, stepped, and randomly irregular terrains, significantly enhancing terrain adaptability and perceptual reliability.
In metal additive manufacturing, quality assessment remains heavily reliant on expert experience, while existing AI-based approaches lack interpretability. Method: This paper proposes the first explainable quality assessment framework integrating vision-language models (VLMs) with domain knowledge. It distills metallurgical expertise from academic literature into a VLM and employs attention mechanisms to enable semantic-level defect reasoning and natural-language explanation generation. Contribution/Results: Evaluated on 24 single-bead laser-wire deposition samples, the framework achieves significant improvements over generic VLMs in both assessment accuracy (+12.7% F1-score) and explanation consistency (+18.3% BLEU-4). By grounding visual reasoning in domain semantics and generating human-readable justifications, it enhances transparency and operational utility—directly addressing industry’s dual requirements for interpretability and practicality.
Existing traffic-scene generation methods face two key bottlenecks: full-sequence denoising compromises online responsiveness, while frame-wise prediction lacks explicit object-state guidance; moreover, open datasets predominantly cover routine behaviors, hindering realistic generation of high-risk corner cases. This paper proposes Nexus, a decoupled diffusion framework that introduces partial noise masking training and noise-aware scheduling—enabling disentangled modeling of sequential coherence and scenario challenge during layout generation. We further design fine-grained tokenized diffusion, independent noise-state modeling for dynamic agents, and closed-loop planning co-optimization. To support rigorous evaluation, we construct the first 540-hour high-risk corner-case simulation dataset. Experiments demonstrate a 40% reduction in displacement error and a 20% improvement in closed-loop planning performance, significantly enhancing realism and safety in complex interactive scenarios—including aggressive cut-ins, emergency braking, and collision avoidance.