MAETrack: Unleashing the Potential of Pretrained Geometric Priors for 3D Single Object Tracking
本文针对3D单目标跟踪中预训练几何先验的迁移问题,提出MAETrack框架,通过分层选择性初始化和几何残差门控方法有效改善了从3D重建到3D跟踪的任务适应。
本文针对3D单目标跟踪中预训练几何先验的迁移问题,提出MAETrack框架,通过分层选择性初始化和几何残差门控方法有效改善了从3D重建到3D跟踪的任务适应。
本文针对WAM中未来想象适应性不足的问题,提出ProWAM模型,通过引入执行进度作为中介表示来优化想象利用。
本文通过多智能体强化学习优化低空流体天线网络的动态重构,以解决动态信道、遮挡和干扰问题,提高系统总速率。
本文通过提出TINA+方法,利用扩散一致的无文本逆向技术来探究被删除概念的视觉知识是否仍然存在于图像生成模型中。
This work addresses the challenge that existing 3D semantic occupancy prediction methods struggle to handle heterogeneous indoor and outdoor scenes within a unified framework. To this end, we introduce the cross-scene 3D semantic occupancy prediction task and propose OccAnyScene, a novel framework built upon a pretrained foundation model. OccAnyScene leverages pixel-aligned frustum feature aggregation and a frustum-parameterized Gaussian decoding mechanism to adaptively reconstruct scenes with varying camera configurations and spatial scales. Our approach is the first to enable a single model to consistently process diverse environments while preserving both metric consistency and scene adaptability. It achieves state-of-the-art performance with mIoU scores of 59.92% on Occ-ScanNet (indoor) and 23.06% on SurroundOcc-nuScenes (outdoor), setting new benchmarks for cross-scene 3D semantic occupancy prediction.
本文针对3D单目标跟踪中预训练几何先验的迁移问题,提出MAETrack框架,通过分层选择性初始化和几何残差门控方法有效改善了从3D重建到3D跟踪的任务适应。
本文针对WAM中未来想象适应性不足的问题,提出ProWAM模型,通过引入执行进度作为中介表示来优化想象利用。
本文通过多智能体强化学习优化低空流体天线网络的动态重构,以解决动态信道、遮挡和干扰问题,提高系统总速率。
本文通过提出TINA+方法,利用扩散一致的无文本逆向技术来探究被删除概念的视觉知识是否仍然存在于图像生成模型中。
This work addresses the challenge that existing 3D semantic occupancy prediction methods struggle to handle heterogeneous indoor and outdoor scenes within a unified framework. To this end, we introduce the cross-scene 3D semantic occupancy prediction task and propose OccAnyScene, a novel framework built upon a pretrained foundation model. OccAnyScene leverages pixel-aligned frustum feature aggregation and a frustum-parameterized Gaussian decoding mechanism to adaptively reconstruct scenes with varying camera configurations and spatial scales. Our approach is the first to enable a single model to consistently process diverse environments while preserving both metric consistency and scene adaptability. It achieves state-of-the-art performance with mIoU scores of 59.92% on Occ-ScanNet (indoor) and 23.06% on SurroundOcc-nuScenes (outdoor), setting new benchmarks for cross-scene 3D semantic occupancy prediction.