Institution profile

Harvard-Smithsonian Center for Astrophysics

Academic institutionnorthamerica · us
Official website
Research library6linked papers
Opportunities0open roles
Selected work

Representative Papers

AstroConcepts: A Large-Scale Multi-Label Classification Corpus for Astrophysics

Apr 02, 2026

This study addresses the severe class imbalance in scientific multi-label text classification caused by the extreme long-tail distribution of domain-specific terms. To this end, the authors construct a large-scale corpus comprising 21,702 astrophysics paper abstracts annotated with 2,367 concepts from the Unified Astronomy Thesaurus. They propose a novel frequency-stratified evaluation strategy and systematically compare the performance of traditional machine learning models, neural networks, and lexically constrained large language models (LLMs). The findings reveal that lexically constrained LLMs achieve performance comparable to specialized models without domain-specific fine-tuning, while domain adaptation significantly improves classification of rare terms. The proposed evaluation framework effectively uncovers model robustness disparities across frequency strata, establishing a strong baseline and a new paradigm for tackling extreme imbalance in scientific text classification tasks.

0 citationsRead paper

Do Lexical and Contextual Coreference Resolution Systems Degrade Differently under Mention Noise? An Empirical Study on Scientific Software Mentions

Apr 02, 2026

This study investigates the robustness of different approaches to cross-document coreference resolution for scientific software mentions under mention-level noise. We systematically evaluate unfine-tuned Fuzzy Matching (FM) against Context-Aware Representations (CAR) using controlled noise injection experiments in the SOMD 2026 shared task, analyzing their performance degradation patterns and inference efficiency under boundary perturbations and mention substitutions. Our work reveals, for the first time, complementary failure characteristics between FM and CAR, and quantifies their computational complexity differences: CAR achieves an F1 score of 0.95–0.96 on the official test set—significantly outperforming FM—and demonstrates greater robustness to boundary noise (F1 drops by only 0.07). Moreover, CAR exhibits near-linear inference complexity, making it more suitable for large-scale deployment.

0 citationsRead paper

Harnessing Self-Supervised Deep Learning and Geostationary Remote Sensing for Advancing Wildfire and Associated Air Quality Monitoring: Improved Smoke and Fire Front Masking using GOES and TEMPO Radiance Data

Oct 10, 2025

Addressing severe smoke/cloud confusion, insufficient real-time capability, and scarce annotated data in wildfire monitoring, this paper proposes a self-supervised deep learning framework that fuses hourly radiometric observations from the GOES-18 and TEMPO geostationary satellites. Without requiring manual annotations, the method leverages temporal consistency modeling and multi-source data协同 reconstruction to achieve pixel-level discrimination and dynamic segmentation of smoke, fire pixels, and clouds. The generated smoke and fire masks exhibit high spatiotemporal coherence and demonstrate strong agreement with multi-source ground and satellite observations across multiple real wildfire events in the western United States. Compared to operational products, the approach achieves significantly improved detection accuracy and reduces response latency to the hourly scale. This work establishes a scalable, annotation-light paradigm for near-real-time wildfire spread tracking and air quality assessment.

0 citationsRead paper
Recent publications

Latest Papers

AstroConcepts: A Large-Scale Multi-Label Classification Corpus for Astrophysics

Apr 02, 2026

This study addresses the severe class imbalance in scientific multi-label text classification caused by the extreme long-tail distribution of domain-specific terms. To this end, the authors construct a large-scale corpus comprising 21,702 astrophysics paper abstracts annotated with 2,367 concepts from the Unified Astronomy Thesaurus. They propose a novel frequency-stratified evaluation strategy and systematically compare the performance of traditional machine learning models, neural networks, and lexically constrained large language models (LLMs). The findings reveal that lexically constrained LLMs achieve performance comparable to specialized models without domain-specific fine-tuning, while domain adaptation significantly improves classification of rare terms. The proposed evaluation framework effectively uncovers model robustness disparities across frequency strata, establishing a strong baseline and a new paradigm for tackling extreme imbalance in scientific text classification tasks.

0 citationsRead paper

Do Lexical and Contextual Coreference Resolution Systems Degrade Differently under Mention Noise? An Empirical Study on Scientific Software Mentions

Apr 02, 2026

This study investigates the robustness of different approaches to cross-document coreference resolution for scientific software mentions under mention-level noise. We systematically evaluate unfine-tuned Fuzzy Matching (FM) against Context-Aware Representations (CAR) using controlled noise injection experiments in the SOMD 2026 shared task, analyzing their performance degradation patterns and inference efficiency under boundary perturbations and mention substitutions. Our work reveals, for the first time, complementary failure characteristics between FM and CAR, and quantifies their computational complexity differences: CAR achieves an F1 score of 0.95–0.96 on the official test set—significantly outperforming FM—and demonstrates greater robustness to boundary noise (F1 drops by only 0.07). Moreover, CAR exhibits near-linear inference complexity, making it more suitable for large-scale deployment.

0 citationsRead paper

Harnessing Self-Supervised Deep Learning and Geostationary Remote Sensing for Advancing Wildfire and Associated Air Quality Monitoring: Improved Smoke and Fire Front Masking using GOES and TEMPO Radiance Data

Oct 10, 2025

Addressing severe smoke/cloud confusion, insufficient real-time capability, and scarce annotated data in wildfire monitoring, this paper proposes a self-supervised deep learning framework that fuses hourly radiometric observations from the GOES-18 and TEMPO geostationary satellites. Without requiring manual annotations, the method leverages temporal consistency modeling and multi-source data协同 reconstruction to achieve pixel-level discrimination and dynamic segmentation of smoke, fire pixels, and clouds. The generated smoke and fire masks exhibit high spatiotemporal coherence and demonstrate strong agreement with multi-source ground and satellite observations across multiple real wildfire events in the western United States. Compared to operational products, the approach achieves significantly improved detection accuracy and reduces response latency to the hourly scale. This work establishes a scalable, annotation-light paradigm for near-real-time wildfire spread tracking and air quality assessment.

0 citationsRead paper