Institution profile

China Academy of Space Technology

Academic institutionasia · cn
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Cognitively Layered Data Synthesis for Domain Adaptation of LLMs to Space Situational Awareness

Mar 10, 2026

This work addresses the challenges of transferring large language models (LLMs) to the domain of space situational awareness (SSA), which stem from task-structure misalignment, lack of higher-order cognitive supervision, and inconsistencies between data and engineering standards. To overcome these issues, the authors propose the BD-FDG framework, which—drawing on Bloom’s taxonomy for the first time in domain-adaptive data generation—constructs a continuous gradient of samples spanning nine question types and six cognitive difficulty levels. Domain knowledge is organized via a knowledge-tree structure, and a multidimensional automated quality assessment pipeline yields a high-quality SSA-SFT dataset comprising 230,000 samples. The resulting SSA-LLM-8B, fine-tuned from Qwen3-8B, achieves a 144% (without chain-of-thought) and 176% (with chain-of-thought) improvement in BLEU-1 on in-domain evaluation, attains an arena win rate of 82.21%, and preserves strong general-purpose capabilities.

0 citationsRead paper
Recent publications

Latest Papers

Cognitively Layered Data Synthesis for Domain Adaptation of LLMs to Space Situational Awareness

Mar 10, 2026

This work addresses the challenges of transferring large language models (LLMs) to the domain of space situational awareness (SSA), which stem from task-structure misalignment, lack of higher-order cognitive supervision, and inconsistencies between data and engineering standards. To overcome these issues, the authors propose the BD-FDG framework, which—drawing on Bloom’s taxonomy for the first time in domain-adaptive data generation—constructs a continuous gradient of samples spanning nine question types and six cognitive difficulty levels. Domain knowledge is organized via a knowledge-tree structure, and a multidimensional automated quality assessment pipeline yields a high-quality SSA-SFT dataset comprising 230,000 samples. The resulting SSA-LLM-8B, fine-tuned from Qwen3-8B, achieves a 144% (without chain-of-thought) and 176% (with chain-of-thought) improvement in BLEU-1 on in-domain evaluation, attains an arena win rate of 82.21%, and preserves strong general-purpose capabilities.

0 citationsRead paper