Diffuse2Seg: Diffusion Models Can Segment Anything Without Supervision
本文提出Diffuse2Seg方法,利用文本到图像扩散模型自动生成多粒度实例掩码,无需监督即可实现开放世界实体分割。
本文提出Diffuse2Seg方法,利用文本到图像扩散模型自动生成多粒度实例掩码,无需监督即可实现开放世界实体分割。
This work addresses the inconsistency between training and inference constraints in robot motion generation by proposing ConFlow, a novel framework that integrates constraint guidance directly into flow matching. ConFlow explicitly models task constraints through differentiable obstacle or cost functions and replaces the standard Gaussian prior with a conditional Gaussian process to enforce trajectory smoothness and boundary conditions. Additionally, it leverages infeasible trajectories as negative supervision signals to enhance constraint adherence. Experimental results demonstrate that ConFlow significantly reduces collision rates and improves trajectory quality in dual-robot navigation tasks, outperforming existing flow matching approaches both with and without inference-time guidance.
This work proposes the first end-to-end fully self-supervised method for joint 3D occupancy and motion estimation in autonomous driving, addressing the reliance on costly human annotations or external supervision. By decoupling static and dynamic signed distance fields and incorporating temporal feature aggregation with a cosine similarity constraint on features, the model implicitly learns scene dynamics without any manual labels. The approach is validated on SemanticKITTI, KITTI-MOT, and nuScenes datasets, demonstrating significant reduction in dependence on annotated data. Furthermore, it introduces a strong self-supervised optical flow cue derived from feature similarity, advancing the state of self-supervised modeling of dynamic 3D scenes.
This work addresses the challenge that deep state-space models often fail to accurately capture the underlying dynamics of observed sequences during training. To overcome this limitation, the authors propose a constrained optimization framework that integrates amortized variational inference with classical Bayesian filtering and smoothing techniques, resulting in an Extended Kalman Variational Autoencoder (EKVAE). By imposing structured constraints, the framework enables a disentangled representation of static and dynamic latent variables. Experimental results demonstrate that EKVAE significantly outperforms existing state-of-the-art models in system identification and long-term prediction tasks, while successfully learning semantically meaningful and disentangled latent state representations.
This work addresses a systematic error in 3D bounding box annotations within dynamic scenes, caused by temporal misalignment among sensor scans, which displaces annotations from the true physical trajectories of objects. We are the first to identify and quantify this error in major autonomous driving datasets and propose an offline optimization method grounded in physical motion models to enforce spatiotemporal consistency across LiDAR time-series data. Evaluated on Argoverse 2, MAN TruckScenes, and a custom dataset, our approach corrects annotation offsets as large as 2.5 meters, improving annotation quality by over 17%. Notably, the impact of this error on benchmark evaluations exceeds the performance gains of typical state-of-the-art methods, underscoring its significant confounding effect on algorithm assessment.
本文提出Diffuse2Seg方法,利用文本到图像扩散模型自动生成多粒度实例掩码,无需监督即可实现开放世界实体分割。
This work addresses the inconsistency between training and inference constraints in robot motion generation by proposing ConFlow, a novel framework that integrates constraint guidance directly into flow matching. ConFlow explicitly models task constraints through differentiable obstacle or cost functions and replaces the standard Gaussian prior with a conditional Gaussian process to enforce trajectory smoothness and boundary conditions. Additionally, it leverages infeasible trajectories as negative supervision signals to enhance constraint adherence. Experimental results demonstrate that ConFlow significantly reduces collision rates and improves trajectory quality in dual-robot navigation tasks, outperforming existing flow matching approaches both with and without inference-time guidance.
This work proposes the first end-to-end fully self-supervised method for joint 3D occupancy and motion estimation in autonomous driving, addressing the reliance on costly human annotations or external supervision. By decoupling static and dynamic signed distance fields and incorporating temporal feature aggregation with a cosine similarity constraint on features, the model implicitly learns scene dynamics without any manual labels. The approach is validated on SemanticKITTI, KITTI-MOT, and nuScenes datasets, demonstrating significant reduction in dependence on annotated data. Furthermore, it introduces a strong self-supervised optical flow cue derived from feature similarity, advancing the state of self-supervised modeling of dynamic 3D scenes.
This work addresses the challenge that deep state-space models often fail to accurately capture the underlying dynamics of observed sequences during training. To overcome this limitation, the authors propose a constrained optimization framework that integrates amortized variational inference with classical Bayesian filtering and smoothing techniques, resulting in an Extended Kalman Variational Autoencoder (EKVAE). By imposing structured constraints, the framework enables a disentangled representation of static and dynamic latent variables. Experimental results demonstrate that EKVAE significantly outperforms existing state-of-the-art models in system identification and long-term prediction tasks, while successfully learning semantically meaningful and disentangled latent state representations.
This work addresses a systematic error in 3D bounding box annotations within dynamic scenes, caused by temporal misalignment among sensor scans, which displaces annotations from the true physical trajectories of objects. We are the first to identify and quantify this error in major autonomous driving datasets and propose an offline optimization method grounded in physical motion models to enforce spatiotemporal consistency across LiDAR time-series data. Evaluated on Argoverse 2, MAN TruckScenes, and a custom dataset, our approach corrects annotation offsets as large as 2.5 meters, improving annotation quality by over 17%. Notably, the impact of this error on benchmark evaluations exceeds the performance gains of typical state-of-the-art methods, underscoring its significant confounding effect on algorithm assessment.