SPEAR NeXT Causal Latent Forecasting Across Multiple Horizons for Spectral Temporal Earth Representation Learning
为了解决地球观测中时间信息学习的问题,SPEAR NeXT模型通过因果掩码Transformer预测多时序潜在状态,利用旋转位置嵌入和年月嵌入来表示相对时间和季节性。
为了解决地球观测中时间信息学习的问题,SPEAR NeXT模型通过因果掩码Transformer预测多时序潜在状态,利用旋转位置嵌入和年月嵌入来表示相对时间和季节性。
研究探讨了地球基础模型中物理类型化和几何感知表示的必要性,通过对比不同方法在多种条件下的表现,验证其是否能带来实质性改进。
This work addresses the challenge of agricultural monitoring under complex temporal, phenological, and climatic dynamics, where existing approaches predominantly rely on optical imagery or multimodal data. The study proposes the first self-supervised learning framework that operates solely on SAR intensity images, introducing an enhanced temporal pretraining task integrated with a tailored masking strategy and curriculum learning to effectively capture phenological features for crop type identification without any optical data. Evaluated on the SICKLE benchmark, the method achieves an IoU of 84.9%, substantially outperforming optical-based baselines by 15.3 percentage points and surpassing current SAR-only approaches by 2.2 percentage points, thereby overcoming a critical bottleneck in all-weather representation learning for agricultural remote sensing.
This study addresses the inefficiency and inaccuracy of manual sugarcane emergence monitoring, which struggles to reliably identify missing-plant areas (“bald patches”). To overcome this, the authors propose an automated pipeline leveraging UAV imagery and YOLOv8-based object detection, enhanced by a novel use of Minimum Spanning Trees (MST) to normalize planting orientation. This approach effectively accommodates diverse field layouts and robustly extracts sugarcane rows. Trained on UAV data from multiple agro-climatic zones, the model converts detection outputs into geospatial point clouds and exports them in Well-Known Text (WKT) format for integration with GIS platforms. The resulting high-resolution emergence maps enable precise replanting guidance, thereby enhancing crop yield, resource-use efficiency, and overall agricultural sustainability.
This work addresses the lack of readable, well-structured, and reusable open-source pretraining frameworks for small language models in education and research by introducing a modular PyTorch-based library. The framework employs composable primitives—such as Block, Residual, Repeat, and Parallel—to ensure alignment between model code and architectural diagrams, enabling seamless transition from pedagogical examples to full-scale pretraining. It integrates streaming data processing, mixed-precision training, callback mechanisms, and single-node multi-GPU support, allowing architecture or component substitution without code modification. Experiments demonstrate that a 348M-parameter model achieves 90.6% weak scaling efficiency across four GPUs, closely matching reference implementations. The project includes 27 preset models, comprehensive documentation, and has received positive early community feedback, effectively bridging teaching, research, and engineering practices.
为了解决地球观测中时间信息学习的问题,SPEAR NeXT模型通过因果掩码Transformer预测多时序潜在状态,利用旋转位置嵌入和年月嵌入来表示相对时间和季节性。
研究探讨了地球基础模型中物理类型化和几何感知表示的必要性,通过对比不同方法在多种条件下的表现,验证其是否能带来实质性改进。
This work addresses the challenge of agricultural monitoring under complex temporal, phenological, and climatic dynamics, where existing approaches predominantly rely on optical imagery or multimodal data. The study proposes the first self-supervised learning framework that operates solely on SAR intensity images, introducing an enhanced temporal pretraining task integrated with a tailored masking strategy and curriculum learning to effectively capture phenological features for crop type identification without any optical data. Evaluated on the SICKLE benchmark, the method achieves an IoU of 84.9%, substantially outperforming optical-based baselines by 15.3 percentage points and surpassing current SAR-only approaches by 2.2 percentage points, thereby overcoming a critical bottleneck in all-weather representation learning for agricultural remote sensing.
This study addresses the inefficiency and inaccuracy of manual sugarcane emergence monitoring, which struggles to reliably identify missing-plant areas (“bald patches”). To overcome this, the authors propose an automated pipeline leveraging UAV imagery and YOLOv8-based object detection, enhanced by a novel use of Minimum Spanning Trees (MST) to normalize planting orientation. This approach effectively accommodates diverse field layouts and robustly extracts sugarcane rows. Trained on UAV data from multiple agro-climatic zones, the model converts detection outputs into geospatial point clouds and exports them in Well-Known Text (WKT) format for integration with GIS platforms. The resulting high-resolution emergence maps enable precise replanting guidance, thereby enhancing crop yield, resource-use efficiency, and overall agricultural sustainability.
This work addresses the lack of readable, well-structured, and reusable open-source pretraining frameworks for small language models in education and research by introducing a modular PyTorch-based library. The framework employs composable primitives—such as Block, Residual, Repeat, and Parallel—to ensure alignment between model code and architectural diagrams, enabling seamless transition from pedagogical examples to full-scale pretraining. It integrates streaming data processing, mixed-precision training, callback mechanisms, and single-node multi-GPU support, allowing architecture or component substitution without code modification. Experiments demonstrate that a 348M-parameter model achieves 90.6% weak scaling efficiency across four GPUs, closely matching reference implementations. The project includes 27 preset models, comprehensive documentation, and has received positive early community feedback, effectively bridging teaching, research, and engineering practices.