Score
Designs methods to estimate 3D structure and canopy height from aerial or remote-sensing data, producing height models, reconstruction pipelines, and validation against ground truth.
This work addresses large-scale, unconstrained 3D reconstruction and novel-view synthesis from sparse, uncalibrated, multi-source heterogeneous images (ground-level, surveillance, aerial) in emergency scenarios such as disaster response and law enforcement—characterized by mixed camera types, extreme illumination and viewpoint variations. Methodologically, we introduce the first real-world, cross-altitude, multi-view, uncalibrated benchmark dataset; propose an end-to-end framework integrating multi-scale NeRF, robust structure-from-motion (SfM), illumination normalization, and cross-modal geometric consistency optimization. Crucially, we decouple evaluation of camera calibration accuracy from rendering quality for the first time, thereby clarifying the true boundaries of practical challenges. Experiments demonstrate that our approach generates high-fidelity, navigable 3D models in complex real-world settings, significantly outperforming existing baselines and advancing prior-free large-scale reconstruction research.
To address the high cost and poor quantifiability of in-field, live-plant 3D modeling and high-throughput phenotyping, this study develops an open-source, low-cost (under USD 15 per plant) vision-only photogrammetric system. Methodologically, it introduces the first lightweight Structure-from-Motion (SfM) pipeline tailored for dynamic field-grown wheat canopies, integrating OpenCV and MeshLab to enable fully automated processing—from multi-view images to point clouds to geometric features. Innovatively, it proposes objective, quantitative metrics for canopy architecture classification (erect vs. prostrate), and successfully extracts 12 key phenotypic traits—including plant height, canopy width, leaf inclination angle, and convex hull volume—with a mean absolute error <3.2% and throughput of 5 minutes per sample. This system establishes a reproducible, accessible 3D phenotyping paradigm for crop science.
Existing 3D reconstruction methods struggle to simultaneously achieve photorealistic appearance, geometric accuracy, and agronomic utility in crop scenes, and lack standardized evaluation benchmarks tailored to repetitive multi-view drone imagery. This work introduces the first public benchmark dataset for precision agriculture, comprising 91 field plots of maize, soybean, wheat, and oat, along with 88,830 high-resolution RGB images. Two evaluation tracks are established: one focusing on optimized scene reconstruction using NeRF and 3D Gaussian Splatting (3DGS), and the other assessing zero-shot geometry estimation by pre-trained feedforward models such as MapAnything. Experiments show that Splatfacto-big achieves the best visual fidelity, while Scaffold-GS excels in depth and canopy height recovery; notably, only MapAnything reliably recovers absolute scale, whereas other feedforward models exhibit significant scale bias.
To address the high cost and low efficiency of diameter-at-breast-height (DBH) measurement in forest inventory, this paper proposes a low-cost, end-to-end automated estimation method using a single consumer-grade 360° camera. The method innovatively integrates Structure-from-Motion (SfM) 3D reconstruction, projection-based Grounded SAM for semantic segmentation, and RANSAC-robust circle fitting to precisely localize tree stem cross-sections and invert DBH directly from spherical imagery. An interactive visualization tool is developed to support manual verification. Evaluated on 43 trees and 61 field-measured samples, the approach achieves a median absolute relative error of 5–9%, with accuracy only 2–4% lower than LiDAR-based methods, while reducing hardware cost by two to three orders of magnitude. This work establishes a scalable, easily deployable paradigm for precise forest monitoring in resource-constrained settings.
Monocular remote sensing image-based 3D building reconstruction is hindered by reliance on task-specific architectures and dense, labor-intensive supervision. Method: This paper presents the first systematic evaluation and adaptation of the general-purpose image-to-3D foundation model SAM 3D to remote sensing. We propose a “segment–reconstruct–compose” pipeline for structured, city-scale 3D modeling and introduce the first SAM extension tailored for remote sensing 3D reconstruction. To objectively assess geometric fidelity, we propose CLIP-based Multi-Modal Distance (CMMD), a novel metric quantifying reconstruction quality. Contribution/Results: Evaluated on the NYC Urban Dataset, our approach significantly outperforms TRELLIS, yielding more coherent roof geometries with sharper boundaries. Both FID and CMMD scores show substantial improvement. This work breaks the dependency on task-specific designs and empirically validates the feasibility and effectiveness of leveraging general-purpose foundation models for large-scale urban 3D scene reconstruction.
Quantifying the impact of climate change on tree primary growth has been hindered by the lack of efficient methods for monitoring whole-canopy branch elongation. This study proposes a high-precision three-dimensional reconstruction approach that integrates low-cost unmanned aerial vehicles with a multi-camera crane-based system (CraneCam), achieving millimeter-scale (5–6 mm) modeling of entire deciduous tree canopies in real-world conditions with 92%–98% reconstruction completeness. The authors introduce 3D-printed reference branches to rigorously evaluate fine-structure reconstruction fidelity and advance whole-tree skeletonization analysis, substantially enhancing both the accuracy and spatial coverage of fine-branch elongation dynamics.
This work addresses the critical limitation of existing large remote sensing models—their inability to perceive height, a key vertical dimension essential for spatial reasoning in complex scenes. To bridge this gap, we propose GeoHeightChat, the first height-aware multimodal remote sensing understanding framework. Our approach introduces a vision-language model–driven data generation pipeline to construct two novel benchmarks, GeoHeight-Bench and GeoHeight-Bench+, and integrates height awareness through prompt engineering, metadata extraction, and implicit injection of geometric features. This enables scalable height annotation and interactive height-based reasoning. Experimental results demonstrate that GeoHeightChat substantially mitigates the “vertical blind spot” of current models, achieving significant performance gains in relative height analysis and terrain-awareness tasks.
This study addresses the critical lack of building height data—missing for over 95% of global structures—which hinders low-altitude aerial systems reliant on preconstructed 3D environments. The authors propose the Location Prior Generation Framework (LPGF), introducing a structured, quality-gated multi-source fusion strategy that integrates Sentinel-2 imagery, UAV telemetry, vehicle GPS trajectories, and OpenStreetMap (OSM) footprints. Building heights are assigned via a three-tier rule hierarchy: OSM tags, floor-count inference, and type-based default values. A Shadow Height Estimation Module (SHEM) is selectively activated only when four stringent quality conditions are met. Without requiring dedicated 3D surveys, LPGF generates reusable urban 3D priors, consistently producing priors for 27 buildings in the MiTra A50 Milan dataset. Manual validation on 15 buildings confirms a mean absolute error of just 3.07 meters for type-default heights—well within the <5-meter uncertainty threshold.
This study addresses the challenge of acquiring accurate dense disparity maps in forestry environments, where intricate branch interlacing and complex canopy structures hinder supervised training of stereo matching networks essential for autonomous drone-based pruning. To overcome this limitation, we present the first high-fidelity synthetic binocular dataset tailored for forest depth estimation, built using Unreal Engine 5. Leveraging 115 high-resolution Quixel Megascans tree models, we simulate imagery matching the ZED Mini stereo camera (baseline: 63 mm, focal length: 2.8 mm) and generate 5,520 rectified stereo image pairs at 1920×1080 resolution across three pitch angles, each accompanied by pixel-accurate disparity ground truth. Combining geometric plausibility, photorealistic appearance, and large-scale precise annotations, this dataset fills a critical gap in supervised training data for forestry applications and will be publicly released to advance related research.
This work addresses the performance degradation in camera localization and 3D reconstruction caused by extreme viewpoint disparities and geometric sparsity across satellite, aerial, and ground-level imagery. To this end, the authors introduce the first benchmark dataset comprising 51 scenes with wide-altitude coverage and orthogonal viewpoints, integrating both real and synthetic images. They further propose SkyNet, a curriculum learning–driven model that progressively learns cross-view correspondences to enhance multi-view geometric consistency. Evaluated on standard metrics, SkyNet outperforms existing methods by 9.6% in RRA@5 and 18.1% in RTA@5, establishing a new baseline for large-scale, multi-altitude 3D scene understanding.