MOOSE-Star: Unlocking Tractable Training for Scientific Discovery by Breaking the Complexity Barrier
This work addresses the challenge of directly training models for scientific hypothesis generation, formalized as $P(\text{hypothesis}|\text{background})$, which is hindered by combinatorial complexity scaling as $O(N^k)$. To overcome this, we propose the MOOSE-Star framework, which leverages probabilistic equation decomposition, motivation-guided hierarchical search, and bounded combinatorial mechanisms to reduce complexity from exponential to logarithmic. This enables, for the first time, end-to-end trainable modeling of the hypothesis generation process. Evaluated on TOMATO-Star—a large-scale dataset of 108,717 decomposed scientific papers—MOOSE-Star breaks through the “complexity wall” inherent in conventional brute-force sampling, facilitating efficient, scalable generation of scientific discoveries and supporting continual test-time expansion.