DroneGround: Open-Vocabulary Drone Payload Characterization Using Synthetic Data and Grounded Vision-Language Models
为解决无人机载荷识别难题,提出DroneGround框架,利用合成数据和视觉-语言模型实现开放词汇载荷分析,提高识别准确性和泛化能力。
为解决无人机载荷识别难题,提出DroneGround框架,利用合成数据和视觉-语言模型实现开放词汇载荷分析,提高识别准确性和泛化能力。
该研究通过引入Semantic Boundary Predictor,在反向去噪过程中进行一次性干预,以解决合成脸部生成中的人口统计学不平衡问题,无需模型重训或架构修改。
This work addresses the challenges of drone detection, including scarce real-world annotated data, high visual variability of drones, and interference from birds and other visual distractors. To this end, the authors propose SimD3, a high-fidelity synthetic dataset built using Unreal Engine 5 that encompasses diverse environments, weather conditions, lighting scenarios, and flight trajectories. SimD3 is the first to explicitly model heterogeneous payload-carrying drones alongside multiple bird species as visual interferents, leveraging a 360-degree six-camera system to generate high-quality training data in complex scenes. The authors train YOLOv5 and an attention-enhanced variant, YOLOv5m+C3b (which replaces the standard C3 module with a C3b block). Experiments demonstrate that models trained on SimD3 significantly improve small-object drone detection, with YOLOv5m+C3b consistently outperforming baseline models across synthetic, mixed, and multiple unseen real-world datasets.
This study addresses the long-standing citation bias in the literature on the first-digit law, which has predominantly credited Benford (1938) while overlooking Newcomb’s earlier contribution (1881). Through bibliometric and citation network analysis, the paper systematically traces historical citation patterns and quantifies, for the first time, the distorting effect of the naming convention “Benford’s Law” on the recognition of scientific priority. The findings reveal that Raimi’s (1976) formalization of the term significantly reinforced Benford’s exclusive attribution, leading high-impact publications to consistently neglect Newcomb’s foundational work. This demonstrates a mechanism of historical bias in academic citation driven by the perceived authority of established nomenclature.
Existing stereo vision methods suffer from severe accuracy degradation in CCTV-based ultra-long-range 3D localization (up to 5 km), primarily due to insufficient modeling of lens distortion—especially in long-baseline bundle adjustment. Method: We propose the first hybrid distortion model tailored for long-range photogrammetry: it extends the classical polynomial model with high-order terms and employs a lightweight neural network to learn residual distortions, thereby balancing physical interpretability and strong nonlinear fitting capability. The model is integrated into an end-to-end pipeline combining bundle adjustment optimization and GIS coordinate transformation. Results: Experiments demonstrate stable, high-precision 3D position estimation within 5 km, significantly outperforming pure geometric or purely learning-based baselines. The method supports real-time GIS visualization, enabling practical deployment in large-scale surveillance systems.
为解决无人机载荷识别难题,提出DroneGround框架,利用合成数据和视觉-语言模型实现开放词汇载荷分析,提高识别准确性和泛化能力。
该研究通过引入Semantic Boundary Predictor,在反向去噪过程中进行一次性干预,以解决合成脸部生成中的人口统计学不平衡问题,无需模型重训或架构修改。
This work addresses the challenges of drone detection, including scarce real-world annotated data, high visual variability of drones, and interference from birds and other visual distractors. To this end, the authors propose SimD3, a high-fidelity synthetic dataset built using Unreal Engine 5 that encompasses diverse environments, weather conditions, lighting scenarios, and flight trajectories. SimD3 is the first to explicitly model heterogeneous payload-carrying drones alongside multiple bird species as visual interferents, leveraging a 360-degree six-camera system to generate high-quality training data in complex scenes. The authors train YOLOv5 and an attention-enhanced variant, YOLOv5m+C3b (which replaces the standard C3 module with a C3b block). Experiments demonstrate that models trained on SimD3 significantly improve small-object drone detection, with YOLOv5m+C3b consistently outperforming baseline models across synthetic, mixed, and multiple unseen real-world datasets.
This study addresses the long-standing citation bias in the literature on the first-digit law, which has predominantly credited Benford (1938) while overlooking Newcomb’s earlier contribution (1881). Through bibliometric and citation network analysis, the paper systematically traces historical citation patterns and quantifies, for the first time, the distorting effect of the naming convention “Benford’s Law” on the recognition of scientific priority. The findings reveal that Raimi’s (1976) formalization of the term significantly reinforced Benford’s exclusive attribution, leading high-impact publications to consistently neglect Newcomb’s foundational work. This demonstrates a mechanism of historical bias in academic citation driven by the perceived authority of established nomenclature.
Existing stereo vision methods suffer from severe accuracy degradation in CCTV-based ultra-long-range 3D localization (up to 5 km), primarily due to insufficient modeling of lens distortion—especially in long-baseline bundle adjustment. Method: We propose the first hybrid distortion model tailored for long-range photogrammetry: it extends the classical polynomial model with high-order terms and employs a lightweight neural network to learn residual distortions, thereby balancing physical interpretability and strong nonlinear fitting capability. The model is integrated into an end-to-end pipeline combining bundle adjustment optimization and GIS coordinate transformation. Results: Experiments demonstrate stable, high-precision 3D position estimation within 5 km, significantly outperforming pure geometric or purely learning-based baselines. The method supports real-time GIS visualization, enabling practical deployment in large-scale surveillance systems.