🤖 AI Summary
研究针对核聚变领域中异构稀疏数据组织问题,通过分析多种传感器数据特性及时间与频率分辨率之间的权衡,提出了一种大规模多模态波动数据表示方法。
📝 Abstract
Training effective foundation models requires massive and organized datasets, yet scientific domains such as nuclear fusion present unique challenges due to largely heterogeneous and sparse data. Here we characterize the data used in developing such a model: with over 20 sensor types spanning 5 orders of magnitude in sampling rate, mixed tensor structures (point measurements, spectrograms, images), and nonstationary physics. We analyze our input complexity and discuss trade-offs between temporal context and frequency resolution. Our analysis provides a template for representing multi-modal fluctuation data at scale, with implications for both multi-modal control systems and nuclear fusion.