🤖 AI Summary
This study addresses the limitation of existing self-supervised learning methods that overlook physical constraints in scientific data by proposing the Imposter discrimination task. This approach forces encoders to identify anomalies through cross-entity feature permutation, thereby capturing inter-feature physical dependencies and providing self-supervision signals grounded in physical coherence. Validation across seven land surface modeling tasks using ERA5-Land data demonstrates that Imposter complements existing objectives, significantly improving performance in climate classification, carbon flux estimation, and runoff prediction. Beyond confirming the efficacy of physics-aware pretraining, this work elucidates the alignment mechanism between downstream tasks and pretraining objectives, establishing a novel paradigm for constructing scientific foundation models.
📝 Abstract
Scientific data often describe entities whose features are jointly governed by the laws of physics, yet existing self-supervised learning (SSL) objectives largely ignore this physical coherence. We introduce imposter, a discriminative pretext task that replaces subsets of an entity's features with real observations donated by another entity and trains the encoder to identify the swapped features. Because every donated value is individually plausible, the task can only be solved by learning cross-feature physical dependencies. We evaluate the proposed objectives on global ERA5-Land reanalysis data using 21 environmental variables and assess the learned representations on seven downstream tasks spanning climate classification, carbon flux estimation, and streamflow prediction. Our study includes, to our knowledge, the first systematic comparison of self-supervised objectives for land-surface modeling under a shared architecture and pre-training budget. We find that the most effective pretext task depends on the downstream task family rather than any single objective's superiority, and that imposter provides complementary information when combined with existing SSL objectives. These results suggest that physical coherence is a valuable new source of self-supervision for scientific foundation models.