Diffuse2Seg: Diffusion Models Can Segment Anything Without Supervision
本文提出Diffuse2Seg方法,利用文本到图像扩散模型自动生成多粒度实例掩码,无需监督即可实现开放世界实体分割。
本文提出Diffuse2Seg方法,利用文本到图像扩散模型自动生成多粒度实例掩码,无需监督即可实现开放世界实体分割。
为解决端到端驾驶系统在策略引发状态下的安全问题,本文提出RoG-DAgger方法,通过短时运动预测构建高质量专家示例并适时调整控制权,以改善模型性能。
Current autonomous driving evaluation methods rely on aggregate metrics that struggle to effectively capture system failures in high-risk, low-frequency scenarios. To address this limitation, this work proposes RISC—a risk-informed evaluation protocol that enables model-agnostic and interpretable stress testing through computable risk slices, lightweight data annotation, and risk-guided sampling. Furthermore, the framework leverages large language models to assist in identifying critical yet often overlooked scenarios. Evaluated on monocular pedestrian perception tasks, RISC dramatically improves the detection rate of critical failures from 34.0% to 98.5%, demonstrating its superior capability in efficiently uncovering high-risk system deficiencies.
This work addresses the computational bottleneck in closed-loop traffic simulation, where efficiently generating multi-agent behaviors that are simultaneously consistent, controllable, and interactive remains challenging—particularly for real-time replanning in autonomous driving. The authors propose a diffusion-based traffic scene generation framework conditioned on instance-level scene context and multimodal behavioral priors, augmented with a test-time guidance mechanism to modulate safety-critical behaviors. Innovatively integrating proposal priors with a compact latent action representation, the method significantly improves sampling efficiency without retraining and enables flexible trade-offs during inference among realism, safety, and controllability. Experiments on the Waymo Open Motion Dataset demonstrate that the approach achieves a strong balance across these desiderata in diverse interactive scenarios while substantially reducing per-step inference latency.
This work addresses the challenge of multi-vehicle cooperative parking, where sub-meter positioning accuracy and real-time collision avoidance impose conflicting demands. To reconcile these objectives, the authors propose CoPark, a multi-agent self-play reinforcement learning framework built upon a residual policy architecture that integrates offline planning with online reactive correction. Specifically, longitudinal maneuvers employ threat-aware action modulation to yield right-of-way, while lateral control adheres to a reference trajectory to ensure precision, complemented by a closed-loop optimization layer that corrects terminal errors. Evaluated on the newly introduced Dragon Lake Parking and DSC3D benchmarks, CoPark achieves zero-shot success rates of 70–85% with collision rates as low as 3–6%, substantially outperforming classical control, imitation learning, and large-scale reinforcement learning baselines, while also exhibiting emergent complex behaviors such as reverse-path yielding.
本文提出Diffuse2Seg方法,利用文本到图像扩散模型自动生成多粒度实例掩码,无需监督即可实现开放世界实体分割。
为解决端到端驾驶系统在策略引发状态下的安全问题,本文提出RoG-DAgger方法,通过短时运动预测构建高质量专家示例并适时调整控制权,以改善模型性能。
Current autonomous driving evaluation methods rely on aggregate metrics that struggle to effectively capture system failures in high-risk, low-frequency scenarios. To address this limitation, this work proposes RISC—a risk-informed evaluation protocol that enables model-agnostic and interpretable stress testing through computable risk slices, lightweight data annotation, and risk-guided sampling. Furthermore, the framework leverages large language models to assist in identifying critical yet often overlooked scenarios. Evaluated on monocular pedestrian perception tasks, RISC dramatically improves the detection rate of critical failures from 34.0% to 98.5%, demonstrating its superior capability in efficiently uncovering high-risk system deficiencies.
This work addresses the computational bottleneck in closed-loop traffic simulation, where efficiently generating multi-agent behaviors that are simultaneously consistent, controllable, and interactive remains challenging—particularly for real-time replanning in autonomous driving. The authors propose a diffusion-based traffic scene generation framework conditioned on instance-level scene context and multimodal behavioral priors, augmented with a test-time guidance mechanism to modulate safety-critical behaviors. Innovatively integrating proposal priors with a compact latent action representation, the method significantly improves sampling efficiency without retraining and enables flexible trade-offs during inference among realism, safety, and controllability. Experiments on the Waymo Open Motion Dataset demonstrate that the approach achieves a strong balance across these desiderata in diverse interactive scenarios while substantially reducing per-step inference latency.
This work addresses the challenge of multi-vehicle cooperative parking, where sub-meter positioning accuracy and real-time collision avoidance impose conflicting demands. To reconcile these objectives, the authors propose CoPark, a multi-agent self-play reinforcement learning framework built upon a residual policy architecture that integrates offline planning with online reactive correction. Specifically, longitudinal maneuvers employ threat-aware action modulation to yield right-of-way, while lateral control adheres to a reference trajectory to ensure precision, complemented by a closed-loop optimization layer that corrects terminal errors. Evaluated on the newly introduced Dragon Lake Parking and DSC3D benchmarks, CoPark achieves zero-shot success rates of 70–85% with collision rates as low as 3–6%, substantially outperforming classical control, imitation learning, and large-scale reinforcement learning baselines, while also exhibiting emergent complex behaviors such as reverse-path yielding.