ReconPlusGen: Injecting Reconstruction Prior into Multi-view 3D Generation through Noise Inversion and Modulation
该研究通过噪声反转和调制,将重建先验注入多视图3D生成中,以完成未观察区域并细化可见几何结构。
该研究通过噪声反转和调制,将重建先验注入多视图3D生成中,以完成未观察区域并细化可见几何结构。
研究解决了机器翻译中目标长度选择问题,提出了一种基于平均预测熵的无训练长度选择器Entropy-Valley(EV),以优化翻译质量。
Existing world models for autonomous driving struggle to balance robustness and interpretability between pixel-level prediction and latent representations. This work proposes a hybrid world modeling framework that unifies pixel supervision and latent representation learning through a two-stage training strategy: in the pretraining stage, it jointly optimizes video latent feature prediction and pixel-level frame reconstruction; in the fine-tuning stage, it relies solely on latent features to drive an action expert. This approach achieves concurrent improvements in fine-grained spatiotemporal reasoning and noise robustness. Evaluated on NAVSIM v1/v2, the method significantly outperforms baselines based exclusively on pixels or latent representations and establishes a new benchmark for evaluating noise robustness in autonomous driving world models.
This work addresses the performance degradation of existing lane perception methods under challenging conditions such as occlusion or missing lane markings, where visual cues are insufficient, and highlights the high cost and poor real-time capability of HD map–dependent approaches. To overcome these limitations, the authors propose a Traffic Flow-aware Module (TFM), which, for the first time, leverages real-time, zero-cost traffic flow information as an auxiliary modality for lane perception—requiring neither additional hardware nor HD maps. TFM employs deep learning to extract dynamic traffic flow features and integrates them with mainstream lane detection models through multimodal fusion. Experiments on the NuScenes and OpenLaneV2 datasets demonstrate consistent performance gains across four state-of-the-art models upon incorporating TFM, with mAP improvements of up to 4.1%, significantly enhancing robustness in complex driving scenarios.
Autonomous trajectory planning in complex urban environments faces key challenges: difficulty in modeling multimodal behavior, poor generalization of single-expert models, and insufficient modeling of vehicle–environment interactions. To address these, this paper proposes an Explicit Mixture-of-Experts (EMoE) dynamic routing framework. Its core contributions are: (1) the first scene-aware explicit MoE architecture, employing a learnable router for task-adaptive expert selection; (2) a multimodal prior query mechanism to enhance diversity-aware trajectory modeling; and (3) an interaction-aware graph neural network coupled with a co-optimization loss function to explicitly capture bidirectional influences between the ego-vehicle and dynamic environmental agents. Evaluated on the NuPlan benchmark, our method achieves state-of-the-art performance across all major test scenarios, significantly improving planning success rate, ride comfort, and safety.
该研究通过噪声反转和调制,将重建先验注入多视图3D生成中,以完成未观察区域并细化可见几何结构。
研究解决了机器翻译中目标长度选择问题,提出了一种基于平均预测熵的无训练长度选择器Entropy-Valley(EV),以优化翻译质量。
Existing world models for autonomous driving struggle to balance robustness and interpretability between pixel-level prediction and latent representations. This work proposes a hybrid world modeling framework that unifies pixel supervision and latent representation learning through a two-stage training strategy: in the pretraining stage, it jointly optimizes video latent feature prediction and pixel-level frame reconstruction; in the fine-tuning stage, it relies solely on latent features to drive an action expert. This approach achieves concurrent improvements in fine-grained spatiotemporal reasoning and noise robustness. Evaluated on NAVSIM v1/v2, the method significantly outperforms baselines based exclusively on pixels or latent representations and establishes a new benchmark for evaluating noise robustness in autonomous driving world models.
This work addresses the performance degradation of existing lane perception methods under challenging conditions such as occlusion or missing lane markings, where visual cues are insufficient, and highlights the high cost and poor real-time capability of HD map–dependent approaches. To overcome these limitations, the authors propose a Traffic Flow-aware Module (TFM), which, for the first time, leverages real-time, zero-cost traffic flow information as an auxiliary modality for lane perception—requiring neither additional hardware nor HD maps. TFM employs deep learning to extract dynamic traffic flow features and integrates them with mainstream lane detection models through multimodal fusion. Experiments on the NuScenes and OpenLaneV2 datasets demonstrate consistent performance gains across four state-of-the-art models upon incorporating TFM, with mAP improvements of up to 4.1%, significantly enhancing robustness in complex driving scenarios.
Autonomous trajectory planning in complex urban environments faces key challenges: difficulty in modeling multimodal behavior, poor generalization of single-expert models, and insufficient modeling of vehicle–environment interactions. To address these, this paper proposes an Explicit Mixture-of-Experts (EMoE) dynamic routing framework. Its core contributions are: (1) the first scene-aware explicit MoE architecture, employing a learnable router for task-adaptive expert selection; (2) a multimodal prior query mechanism to enhance diversity-aware trajectory modeling; and (3) an interaction-aware graph neural network coupled with a co-optimization loss function to explicitly capture bidirectional influences between the ego-vehicle and dynamic environmental agents. Evaluated on the NuPlan benchmark, our method achieves state-of-the-art performance across all major test scenarios, significantly improving planning success rate, ride comfort, and safety.