What Should a Large Language Model See? Physical Invariants as a Data Representation for PDE Discovery

📅 2026-08-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过引入数据解释阶段,使大型语言模型能够像理论学家一样解读场数据,从而提高从实验数据中自动构建场理论的准确性。
📝 Abstract
Understanding how molecular interactions govern macroscopic behaviour is a central challenge in molecular sciences. However, conventional theory building cannot keep pace with the vast datasets modern experimentation routinely produces. Large language models offer a promising route to automating theory construction, but a spatiotemporal field cannot be directly placed in a prompt. Existing models generally learn about the data only through a score measuring how well each proposal fits it. Here we introduce data interpretation, a stage that measures the field into the quantities a theorist would consult and supplies them to the model as a direct input. On a benchmark of simulated fields, interpretation nearly triples the accuracy of recovered equations relative to showing the raw data, at negligible computational cost and without any training. By allowing a language model to read field data as a theorist does, data interpretation offers a practical route to automated field theory construction that can coevolve with experimentation.
Problem

Research questions and friction points this paper is trying to address.

molecular interactions
macroscopic behaviour
theory building
large datasets
experimentation
Innovation

Methods, ideas, or system contributions that make the work stand out.

data interpretation
large language models
PDE discovery
theory construction
spatiotemporal field
🔎 Similar Papers
No similar papers found.