Institution profile

MiroMind

Research institution
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

MOOSE-Star: Unlocking Tractable Training for Scientific Discovery by Breaking the Complexity Barrier

Mar 04, 2026

This work addresses the challenge of directly training models for scientific hypothesis generation, formalized as $P(\text{hypothesis}|\text{background})$, which is hindered by combinatorial complexity scaling as $O(N^k)$. To overcome this, we propose the MOOSE-Star framework, which leverages probabilistic equation decomposition, motivation-guided hierarchical search, and bounded combinatorial mechanisms to reduce complexity from exponential to logarithmic. This enables, for the first time, end-to-end trainable modeling of the hypothesis generation process. Evaluated on TOMATO-Star—a large-scale dataset of 108,717 decomposed scientific papers—MOOSE-Star breaks through the “complexity wall” inherent in conventional brute-force sampling, facilitating efficient, scalable generation of scientific discoveries and supporting continual test-time expansion.

0 citationsRead paper

Multi-Agent Tool-Integrated Policy Optimization

Oct 06, 2025

To address context-length limitations and noisy tool responses in multi-turn tool-integration tasks using large language models (LLMs), this paper proposes a lightweight multi-role collaboration framework that dynamically partitions a single LLM instance into distinct “Planner” and “Executor” roles—eliminating the memory overhead of multi-model deployment. We introduce a novel cross-role credit assignment mechanism, integrating role-specific prompting with reinforcement learning for end-to-end post-training, enabling joint optimization of planning and execution. Evaluated on GAIA-text, WebWalkerQA, and FRAMES benchmarks, our method achieves an average relative performance gain of 18.38%, significantly improving robustness to noisy tool outputs. To the best of our knowledge, this is the first work to realize efficient, end-to-end reinforcement training of multiple collaborative roles within a single LLM.

0 citationsRead paper

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization

Jul 19, 2025

Open-source mathematical reasoning models suffer from low transparency, poor reproducibility, and suboptimal performance. Method: We introduce MiroMind-M1, a fully open-source family of reasoning language models built upon Qwen-2.5, trained via a two-stage reproducible pipeline: (i) supervised fine-tuning (SFT) on 719K math problems augmented with verified chain-of-thought rationales; followed by (ii) RLVR-based reinforcement learning on a curated 62K high-difficulty problem set. We propose a context-aware multi-stage policy optimization algorithm incorporating progressive sequence-length expansion and adaptive repetition penalty to enhance RL stability and token efficiency. Contribution/Results: MiroMind-M1 achieves state-of-the-art performance among open-source models of comparable scale (7B/32B) on AIME24, AIME25, and MATH benchmarks. Crucially, we fully release model weights, training data, and configuration scripts—enabling transparent, reproducible mathematical reasoning research.

0 citationsRead paper
Recent publications

Latest Papers

MOOSE-Star: Unlocking Tractable Training for Scientific Discovery by Breaking the Complexity Barrier

Mar 04, 2026

This work addresses the challenge of directly training models for scientific hypothesis generation, formalized as $P(\text{hypothesis}|\text{background})$, which is hindered by combinatorial complexity scaling as $O(N^k)$. To overcome this, we propose the MOOSE-Star framework, which leverages probabilistic equation decomposition, motivation-guided hierarchical search, and bounded combinatorial mechanisms to reduce complexity from exponential to logarithmic. This enables, for the first time, end-to-end trainable modeling of the hypothesis generation process. Evaluated on TOMATO-Star—a large-scale dataset of 108,717 decomposed scientific papers—MOOSE-Star breaks through the “complexity wall” inherent in conventional brute-force sampling, facilitating efficient, scalable generation of scientific discoveries and supporting continual test-time expansion.

0 citationsRead paper

Multi-Agent Tool-Integrated Policy Optimization

Oct 06, 2025

To address context-length limitations and noisy tool responses in multi-turn tool-integration tasks using large language models (LLMs), this paper proposes a lightweight multi-role collaboration framework that dynamically partitions a single LLM instance into distinct “Planner” and “Executor” roles—eliminating the memory overhead of multi-model deployment. We introduce a novel cross-role credit assignment mechanism, integrating role-specific prompting with reinforcement learning for end-to-end post-training, enabling joint optimization of planning and execution. Evaluated on GAIA-text, WebWalkerQA, and FRAMES benchmarks, our method achieves an average relative performance gain of 18.38%, significantly improving robustness to noisy tool outputs. To the best of our knowledge, this is the first work to realize efficient, end-to-end reinforcement training of multiple collaborative roles within a single LLM.

0 citationsRead paper

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization

Jul 19, 2025

Open-source mathematical reasoning models suffer from low transparency, poor reproducibility, and suboptimal performance. Method: We introduce MiroMind-M1, a fully open-source family of reasoning language models built upon Qwen-2.5, trained via a two-stage reproducible pipeline: (i) supervised fine-tuning (SFT) on 719K math problems augmented with verified chain-of-thought rationales; followed by (ii) RLVR-based reinforcement learning on a curated 62K high-difficulty problem set. We propose a context-aware multi-stage policy optimization algorithm incorporating progressive sequence-length expansion and adaptive repetition penalty to enhance RL stability and token efficiency. Contribution/Results: MiroMind-M1 achieves state-of-the-art performance among open-source models of comparable scale (7B/32B) on AIME24, AIME25, and MATH benchmarks. Crucially, we fully release model weights, training data, and configuration scripts—enabling transparent, reproducible mathematical reasoning research.

0 citationsRead paper