Scientific Image Synthesis: Benchmarking, Methodologies, and Downstream Utility

📅 2026-01-17
🏛️ arXiv.org
📈 Citations: 7
Influential: 0
📄 PDF
🤖 AI Summary
研究通过提出逻辑驱动框架ImgCoder和评估基准SciGenBench,解决科学图像合成中的视觉-逻辑不一致问题,提高下游推理能力。
📝 Abstract
While synthetic data has proven effective for improving scientific reasoning in the text domain, multimodal reasoning remains constrained by the difficulty of synthesizing scientifically rigorous images. Existing Text-to-Image (T2I) models often produce outputs that are visually plausible yet scientifically incorrect, resulting in a persistent visual-logic divergence that limits their value for downstream reasoning. Motivated by recent advances in next-generation T2I models, we conduct a systematic study of scientific image synthesis across generation paradigms, evaluation, and downstream use. We analyze both direct pixel-based generation and programmatic synthesis, and propose ImgCoder, a logic-driven framework that follows an explicit"understand - plan - code"workflow to improve structural precision. To rigorously assess scientific correctness, we introduce SciGenBench, which evaluates generated images based on information utility and logical validity. Our evaluation reveals systematic failure modes in pixel-based models and highlights a fundamental expressiveness-precision trade-off. Finally, we show that fine-tuning Large Multimodal Models (LMMs) on rigorously verified synthetic scientific images yields consistent reasoning gains, with potential scaling trends analogous to the text domain, validating high-fidelity scientific synthesis as a viable path to unlocking massive multimodal reasoning capabilities.
Problem

Research questions and friction points this paper is trying to address.

scientific image synthesis
visual-logic divergence
downstream reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

ImgCoder
SciGenBench
scientific correctness
Large Multimodal Models (LMMs)
multimodal reasoning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Honglin Lin
Honglin Lin
SJTU
Z
Zheng Liu
Peking University, OpenDataLab, Shanghai Artificial Intelligence Laboratory
C
Chonghan Qin
The University of Hong Kong, OpenDataLab, Shanghai Artificial Intelligence Laboratory
Qizhi Pei
Qizhi Pei
PhD Student, Gaoling School of Artificial Intelligence, Renmin University of China
LLMData SynthesisAI4Science
Y
Yu Li
OpenDataLab, Shanghai Artificial Intelligence Laboratory
Z
Zhanping Zhong
Shanghai Jiao Tong University, OpenDataLab, Shanghai Artificial Intelligence Laboratory
Xin Gao
Xin Gao
Shanghai AI Laboratory & SJTU
ML、NLP、LLM
Yanfeng Wang
Yanfeng Wang
Shanghai Jiao Tong University
Conghui He
Conghui He
Shanghai AI Laboratory
Data-centric AILLMDocument Intelligence
Lijun Wu
Lijun Wu
Shanghai AI Laboratory
MLLLMAI4Science