Rethinking Domain Specialization for Open-Ended Scientific Reasoning in Astronomy Language Models

📅 2026-09-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究探讨了在开放性科学推理中,何时需要领域特定微调,并通过天文学问题集对比通用和领域特化模型的效果。
📝 Abstract
Domain-specialized language models are widely used for scientific question answering, but stronger general-purpose systems raise a sharper question: when does domain-specific fine-tuning remain valuable for open-ended scientific reasoning? We study this in astronomy with a curated QA benchmark from publicly available 2017--2026 Olympiad-style materials. The free-response subset contains 300 questions, including 204 text-only and 96 image-linked examples. We compare open-weight and API-served general-purpose, multimodal, and astronomy-specialized models using judge-based correctness and complementary reference metrics. Strong general-purpose models establish the highest correctness baseline in this testbed, while analyses of metric agreement, judge sensitivity, benchmark composition, and modality reveal variation not captured by a single leaderboard. These results motivate treating domain specialization as a task- and deployment-dependent property and highlight the role of domain-specific evaluation in determining which models, capabilities, and evaluation criteria are appropriate for scientific workflows.
Problem

Research questions and friction points this paper is trying to address.

Domain Specialization
Open-Ended Scientific Reasoning
Astronomy Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

domain specialization
open-ended scientific reasoning
astronomy language models
🔎 Similar Papers
No similar papers found.
V
Vanessa Lama
Oak Ridge National Laboratory, Oak Ridge, Tennessee, USA
Sanjay Das
Sanjay Das
University of Texas at Dallas
Deep learningHardware AcceleratorsHardware testing & securityFunctional safety
E
Emily Herron
Oak Ridge National Laboratory, Oak Ridge, Tennessee, USA
Y
Yuan-Sen Ting
The Ohio State University, Columbus, Ohio, USA; Max-Planck-Institut für Astronomie, Königstuhl 17, D-69117 Heidelberg, Germany
T
Tijmen de Haan
QUP/IPNS, High Energy Accelerator Research Organization (KEK), Tsukuba, Ibaraki, Japan
Junqi Yin
Junqi Yin
National Center for Computational Sciences, Oak Ridge National Laboratory
HPCAIFoundation ModelMaterials Science
Tirthankar Ghosal
Tirthankar Ghosal
Oak Ridge National Laboratory
Natural Language ProcessingMachine LearningArtificial IntelligenceInformation Extraction
Feiyi Wang
Feiyi Wang
Distinguished Research Scientist & Group Leader, Analytics and AI Methods at Scale, NCCS/ORNL
HPCAI for Science at Scale