OmniScientist: An Omni-Modal Omni-Discipline AI Scientist

📅 2026-08-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Current AI research systems struggle to harness the spatiotemporal and cross-modal relationships embedded in raw multimodal data—such as images, signals, and videos—limiting the completeness of scientific discovery. This work proposes OmniScientist, the first end-to-end, fully multimodal AI scientist, which integrates a multimodal perception layer with three autonomous agents for ideation, experimentation, and writing within a deterministic scientific pipeline to drive interdisciplinary research directly from heterogeneous raw evidence. The system incorporates a full-cycle perception mechanism and code-based automated validation to ensure novelty, statistical validity, and traceability. Evaluated across 36 real-world cases spanning five disciplines and four evidence types, OmniScientist autonomously generated complete scientific papers from raw data, achieving an average score of 6.3 and significantly outperforming baseline systems that rely solely on scalar features in seven evaluation dimensions and 85% of pairwise comparisons.
📝 Abstract
Recent advances in foundation models have enabled AI scientists to automate increasingly complete research workflows, from hypothesis generation and code execution to manuscript preparation. Yet workflow coverage alone does not provide access to the full evidence on which scientific discovery depends. Existing systems typically reason over text, code, labels, or precomputed summaries, leaving scientifically decisive spatial, temporal, cross-channel, and procedural relations unavailable to the agent. We introduce OmniScientist, an end-to-end, omni-modal AI scientist that conducts multidisciplinary research directly from heterogeneous raw evidence. A perception layer and 3 autonomous agents for ideation, experiment, and writeup operate within a deterministic pipeline, allowing observations to shape research questions, experimental decisions, and final claims throughout the research lifecycle. By running idea, rigour, and claim checks in code, the system enforces novelty screening, statistical validity, execution provenance, and numerical traceability. We evaluate OmniScientist on 36 real-data cases spanning 5 discipline families, 4 families of scientific evidence, and modalities including images, signals, audio, video, 3-D structures, trajectories, tables, formulae, and graphs. The system completes the full path from raw data to a compiled manuscript in all 36 cases and achieves a mean overall paper score of 6.3 with the reference reasoning backbone. In paired comparisons against a blind variant that receives only precomputed scalar features, direct perception improves all 7 evaluation dimensions and wins 85% of head-to-head judgments. These results show that lifecycle-wide perception is essential for evidence-grounded scientific discovery and provides a practical path toward broadly capable AI scientists.
Problem

Research questions and friction points this paper is trying to address.

scientific discovery
omni-modal perception
raw evidence
multidisciplinary research
AI scientist
Innovation

Methods, ideas, or system contributions that make the work stand out.

omni-modal perception
AI scientist
end-to-end scientific discovery
autonomous research agents
evidence-grounded reasoning
🔎 Similar Papers
No similar papers found.