FFVO: A Feedforward Pose Decoder for Long-Horizon Visual Odometry
为解决长视频中姿态估计的计算成本、上下文模糊性和时间不稳定性问题,提出FFVO方法,通过紧凑表示、层次解码器及轨迹监督实现高效稳定姿态估计。
为解决长视频中姿态估计的计算成本、上下文模糊性和时间不稳定性问题,提出FFVO方法,通过紧凑表示、层次解码器及轨迹监督实现高效稳定姿态估计。
This work addresses the lack of natural dynamics in static, large-scale 3D Gaussian splatting scenes by proposing a time-conditioned deformation field method that integrates a pretrained video diffusion model. Through an iterative data–model update and render–optimization pipeline, the approach injects environmental dynamics—such as swaying vegetation—while preserving rigid structures. It represents the first effort to incorporate diffusion priors into scene-level 3D animation generation, overcoming the limitations of existing methods that are confined to small regions or object-centric settings. The framework enables modeling of distributed, subtle motions across complex outdoor environments. Experiments on five real-world, large-scale outdoor scenes demonstrate that the method produces high-quality novel-view animations, significantly enhancing visual realism and viewer immersion.
This work addresses the limitations of traditional sampling-based motion planning algorithms in real-time performance and integration with modern AI research workflows by presenting a systematic upgrade to the open-source OMPL library. For the first time, hardware acceleration support—encompassing GPUs and FPGAs—is introduced into OMPL, alongside enhanced compatibility with mainstream AI toolchains. The extension supports diverse planning paradigms, including asymptotically optimal planning, lazy sampling, constraint handling, and task specifications expressed in linear temporal logic. These advancements substantially improve computational efficiency and scalability, reinforcing OMPL’s foundational role in motion planning while significantly broadening its applicability to complex intelligent systems and autonomous robotic platforms.
This work addresses limitations in existing model merging approaches, which either neglect the geometric structure of the loss landscape or rely on computationally expensive full-space Hessian approximations, thereby constraining effective knowledge integration. The authors formulate model merging as computing the Fréchet mean on a Riemannian manifold within the low-rank subspace spanned by task vectors, employing the expected Hessian as the metric. This formulation establishes, for the first time, a theoretical link between local curvature and epistemic uncertainty. A rigorous error bound for the merged model is derived, and curvature-aware and spectral methods are shown to be special cases of this unified framework. Experiments on eight image classification tasks using fine-tuned CLIP-ViT models demonstrate that the proposed method consistently outperforms existing baselines in both average and worst-case cross-task accuracy across all backbone architectures.
This work addresses the limitation of current autonomous driving systems, which struggle to leverage the vast amounts of unstructured dashcam video due to a lack of large-scale, structured multimodal sensor data. The authors propose a generative modeling paradigm that, for the first time, converts monocular dashcam footage into high-fidelity multimodal autonomous vehicle logs—comprising multi-view images and LiDAR point clouds—without requiring paired training data. Their approach integrates 4D Gaussian splatting for novel view synthesis and pseudo-paired data construction, coupled with diffusion models to enable high-quality cross-modal generation. The resulting synthetic data demonstrates exceptional fidelity and realism, effectively transforming long-tail driving scenarios from internet-sourced videos into standardized multimodal formats suitable for training and validation.
为解决长视频中姿态估计的计算成本、上下文模糊性和时间不稳定性问题,提出FFVO方法,通过紧凑表示、层次解码器及轨迹监督实现高效稳定姿态估计。
This work addresses the lack of natural dynamics in static, large-scale 3D Gaussian splatting scenes by proposing a time-conditioned deformation field method that integrates a pretrained video diffusion model. Through an iterative data–model update and render–optimization pipeline, the approach injects environmental dynamics—such as swaying vegetation—while preserving rigid structures. It represents the first effort to incorporate diffusion priors into scene-level 3D animation generation, overcoming the limitations of existing methods that are confined to small regions or object-centric settings. The framework enables modeling of distributed, subtle motions across complex outdoor environments. Experiments on five real-world, large-scale outdoor scenes demonstrate that the method produces high-quality novel-view animations, significantly enhancing visual realism and viewer immersion.
This work addresses the limitations of traditional sampling-based motion planning algorithms in real-time performance and integration with modern AI research workflows by presenting a systematic upgrade to the open-source OMPL library. For the first time, hardware acceleration support—encompassing GPUs and FPGAs—is introduced into OMPL, alongside enhanced compatibility with mainstream AI toolchains. The extension supports diverse planning paradigms, including asymptotically optimal planning, lazy sampling, constraint handling, and task specifications expressed in linear temporal logic. These advancements substantially improve computational efficiency and scalability, reinforcing OMPL’s foundational role in motion planning while significantly broadening its applicability to complex intelligent systems and autonomous robotic platforms.
This work addresses limitations in existing model merging approaches, which either neglect the geometric structure of the loss landscape or rely on computationally expensive full-space Hessian approximations, thereby constraining effective knowledge integration. The authors formulate model merging as computing the Fréchet mean on a Riemannian manifold within the low-rank subspace spanned by task vectors, employing the expected Hessian as the metric. This formulation establishes, for the first time, a theoretical link between local curvature and epistemic uncertainty. A rigorous error bound for the merged model is derived, and curvature-aware and spectral methods are shown to be special cases of this unified framework. Experiments on eight image classification tasks using fine-tuned CLIP-ViT models demonstrate that the proposed method consistently outperforms existing baselines in both average and worst-case cross-task accuracy across all backbone architectures.
This work addresses the limitation of current autonomous driving systems, which struggle to leverage the vast amounts of unstructured dashcam video due to a lack of large-scale, structured multimodal sensor data. The authors propose a generative modeling paradigm that, for the first time, converts monocular dashcam footage into high-fidelity multimodal autonomous vehicle logs—comprising multi-view images and LiDAR point clouds—without requiring paired training data. Their approach integrates 4D Gaussian splatting for novel view synthesis and pseudo-paired data construction, coupled with diffusion models to enable high-quality cross-modal generation. The resulting synthetic data demonstrates exceptional fidelity and realism, effectively transforming long-tail driving scenarios from internet-sourced videos into standardized multimodal formats suitable for training and validation.