MeRoPE: Metric Rotary Position Embedding for Camera-Controlled Video Generation
本文针对相机控制视频生成中的尺度依赖问题,提出MeRoPE方法,通过保持特征范数和限制注意力分数来提高生成视频与条件姿态的一致性。
本文针对相机控制视频生成中的尺度依赖问题,提出MeRoPE方法,通过保持特征范数和限制注意力分数来提高生成视频与条件姿态的一致性。
In high-density urban scenarios, 3D pedestrian perception remains challenging due to severe occlusion and clutter, while manual annotation of ground truth trajectories is prohibitively expensive—especially for tail-case pedestrians. Method: We introduce the first multi-view LiDAR-camera fusion multi-object tracking benchmark specifically designed for crowded pedestrians, coupled with an offline automatic annotation system that generates trajectory-level ground truth via cross-modal point cloud–image joint reconstruction. Our tracking-by-detection framework employs a density-aware and relation-aware high-resolution representation learning mechanism, jointly modeling pedestrian density distributions and interaction graphs from multi-view images and sparse LiDAR point clouds. Contribution/Results: Evaluated on our newly established benchmark, our method achieves significant improvements in 3D tracking accuracy; the automatic annotation pipeline accelerates labeling efficiency by over 3×. Both code and dataset will be publicly released.
本文针对相机控制视频生成中的尺度依赖问题,提出MeRoPE方法,通过保持特征范数和限制注意力分数来提高生成视频与条件姿态的一致性。
In high-density urban scenarios, 3D pedestrian perception remains challenging due to severe occlusion and clutter, while manual annotation of ground truth trajectories is prohibitively expensive—especially for tail-case pedestrians. Method: We introduce the first multi-view LiDAR-camera fusion multi-object tracking benchmark specifically designed for crowded pedestrians, coupled with an offline automatic annotation system that generates trajectory-level ground truth via cross-modal point cloud–image joint reconstruction. Our tracking-by-detection framework employs a density-aware and relation-aware high-resolution representation learning mechanism, jointly modeling pedestrian density distributions and interaction graphs from multi-view images and sparse LiDAR point clouds. Contribution/Results: Evaluated on our newly established benchmark, our method achieves significant improvements in 3D tracking accuracy; the automatic annotation pipeline accelerates labeling efficiency by over 3×. Both code and dataset will be publicly released.