Towards Physics-Faithful Generation of Scientific Diagrams

📅 2026-08-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Current text-to-image generation models often produce scientifically plausible-looking illustrations that contain physical inaccuracies, undermining their utility in research and education. To address this, this work proposes Princigram, a novel system that introduces the Structured Physical Chain-of-Thought (SP-CoT) framework. SP-CoT decomposes scientific diagram generation into interpretable reasoning steps—such as scene recognition, process analysis, and law application—explicitly separating visual facts from physical inferences and incorporating symbolic mathematical expressions. Leveraging SP-CoT, the authors construct a large-scale, structurally annotated dataset and introduce VeriphyT2IBench, a new evaluation benchmark enabling fine-grained assessment of physical correctness. Experiments demonstrate that Princigram significantly improves the physical fidelity of generated images on both the GenExam physics subset and VeriphyT2IBench, achieving decomposable and interpretable scientific illustration generation.
📝 Abstract
Text-to-image generation has reached photorealistic quality, yet state-of-the-art systems remain unreliable at producing scientific diagrams, whose value depends not on appearance but on physical faithfulness: correct force directions, valid coordinate systems, consistent thermodynamic states, and equations matching the depicted scenario. Trained on web imagery with physically shallow captions, generic models produce diagrams that look plausible but are physically wrong, harmful in education and scientific communication. We present Princigram, a physics-faithful scientific-diagram generator, and its data pipeline. Our central advance is Structured Physical Chain-of-Thought (SP-CoT): a per-subdiscipline schema that decomposes a physics diagram into an explicit multi-step reasoning chain across six subdisciplines, from scene identification through force or process analysis to governing laws and synthesis. Unlike free-form chain-of-thought, SP-CoT follows a fixed schema with strict fidelity rules that separate visually grounded facts from physically inferred reasoning and type all mathematics symbolically; it serves both as dense training supervision and, at inference, as a structured "thinking" prompt. With it we curate and structurally annotate 4.3 million physics images, of which 115,037 carry expert-level annotation, and adapt a unified multimodal backbone. We further introduce VeriphyT2IBench, whose questions are derived from each held-out diagram's own structured annotation: each diagram becomes an item-specific bank of binary questions about its objects, forces, and states, so a judge model's score decomposes into named physical facts rather than one holistic number. On the physics subset of GenExam and on VeriphyT2IBench, Princigram shows that explicit physics-structured supervision improves the physical faithfulness of generated scientific diagrams.
Problem

Research questions and friction points this paper is trying to address.

scientific diagrams
physical faithfulness
text-to-image generation
physics accuracy
educational reliability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Structured Physical Chain-of-Thought
physics-faithful generation
scientific diagram synthesis
multimodal reasoning
VeriphyT2IBench
🔎 Similar Papers
No similar papers found.