A deep dictionary network-based foundation model for ultra-low-dose CT denoising
为解决超低剂量CT图像噪声问题,提出基于深度字典网络的统一多器官去噪基础模型,通过预训练和微调实现跨区域去噪。
为解决超低剂量CT图像噪声问题,提出基于深度字典网络的统一多器官去噪基础模型,通过预训练和微调实现跨区域去噪。
本文针对固定视角路边感知问题,构建了首个真实世界的基础设施侧语义占用基准InfraOcc,并提出ProSD-Occ方法,通过静态到动态的逐步推理解决该问题。
为解决自动驾驶中长期规划的连续性和一致性问题,提出MomADv2框架,通过选择性状态空间记忆和轨迹残差修正方法提升决策可靠性。
This study addresses the challenges of unreliable historical states and motion evolution in long-horizon planning for end-to-end autonomous driving by proposing StableDrive. The method leverages Mamba operators to construct a selective momentum memory that enhances the robustness of historical representations, while introducing a motion-stage training scaffold to guide the model in perceiving dynamic evolution, thereby enabling efficient single-model deployment without ensembling. Experiments demonstrate that StableDrive achieves state-of-the-art performance on benchmarks such as nuScenes, reducing collision rates by 23.3% and attaining the highest EPDMS score on NAVSIM v2. These results indicate significant improvements in both safety and temporal consistency for long-horizon planning, validating the effectiveness of integrating structured memory mechanisms with stage-aware training in complex driving scenarios.
This work systematically identifies three distinct modes of representation collapse in JEPA-based world models—physical invariance, identifiability, and counterfactual dynamics—even when global latent collapse is avoided. To address these failure modes, the paper introduces PhyLatent, a novel training objective that jointly optimizes dynamics-relevant representations through physical state anchoring, future representation alignment, static visual invariance constraints, counterfactual branch disentanglement, and latent denoising. Moving beyond reliance on global non-collapse assumptions alone, PhyLatent significantly reduces the three collapse rates to 7.53%, 0.95%, and 4.62% on OGBench-Cube, yielding a model-predictive control (MPC) success rate of 78.1%. It further achieves a 98.0% success rate on the TwoRooms task and maintains state-of-the-art performance on Reacher and PushT benchmarks.
为解决超低剂量CT图像噪声问题,提出基于深度字典网络的统一多器官去噪基础模型,通过预训练和微调实现跨区域去噪。
本文针对固定视角路边感知问题,构建了首个真实世界的基础设施侧语义占用基准InfraOcc,并提出ProSD-Occ方法,通过静态到动态的逐步推理解决该问题。
为解决自动驾驶中长期规划的连续性和一致性问题,提出MomADv2框架,通过选择性状态空间记忆和轨迹残差修正方法提升决策可靠性。
This study addresses the challenges of unreliable historical states and motion evolution in long-horizon planning for end-to-end autonomous driving by proposing StableDrive. The method leverages Mamba operators to construct a selective momentum memory that enhances the robustness of historical representations, while introducing a motion-stage training scaffold to guide the model in perceiving dynamic evolution, thereby enabling efficient single-model deployment without ensembling. Experiments demonstrate that StableDrive achieves state-of-the-art performance on benchmarks such as nuScenes, reducing collision rates by 23.3% and attaining the highest EPDMS score on NAVSIM v2. These results indicate significant improvements in both safety and temporal consistency for long-horizon planning, validating the effectiveness of integrating structured memory mechanisms with stage-aware training in complex driving scenarios.
This work systematically identifies three distinct modes of representation collapse in JEPA-based world models—physical invariance, identifiability, and counterfactual dynamics—even when global latent collapse is avoided. To address these failure modes, the paper introduces PhyLatent, a novel training objective that jointly optimizes dynamics-relevant representations through physical state anchoring, future representation alignment, static visual invariance constraints, counterfactual branch disentanglement, and latent denoising. Moving beyond reliance on global non-collapse assumptions alone, PhyLatent significantly reduces the three collapse rates to 7.53%, 0.95%, and 4.62% on OGBench-Cube, yielding a model-predictive control (MPC) success rate of 78.1%. It further achieves a 98.0% success rate on the TwoRooms task and maintains state-of-the-art performance on Reacher and PushT benchmarks.