Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts
本文提出了一种计算高效的两步超参数转移框架,通过跨模型宽度和令牌维度转移学习率,解决大规模混合专家模型的优化问题。
本文提出了一种计算高效的两步超参数转移框架,通过跨模型宽度和令牌维度转移学习率,解决大规模混合专家模型的优化问题。
This work addresses the challenge in reinforcement learning for instruction following where static prompts mismatch the policy’s evolving capabilities, resulting in poorly discriminative reward signals. To resolve this, the authors propose the LLM-as-a-Tutor framework, which treats prompt adaptation as a policy-aware process. Leveraging a large language model as both examiner and tutor, the method identifies non-challenging prompts through pairwise comparisons and monotonically increases task difficulty by appending atomic constraints—ensuring training signals remain calibrated to the policy’s current proficiency. Notably, this approach requires no external scheduling mechanism and significantly outperforms policy-agnostic baselines and existing adaptive methods across three complex instruction-following benchmarks.
This work addresses a critical limitation in existing post-training methods for medical multimodal large language models, which overly prioritize final answer correctness while neglecting the optimization of intermediate reasoning steps, thereby suffering from cascading failures triggered by early errors. To mitigate this, we propose Medical Reasoning-aware Policy Optimization (MRPO), the first reinforcement learning framework tailored for medical multimodal reasoning that incorporates step-aware feedback. MRPO introduces step-level process rewards, applying exponentially stronger penalties to ineffective early reasoning even when the final answer is incorrect, thus effectively interrupting error propagation. Experimental results demonstrate that MRPO consistently outperforms standard GRPO and state-of-the-art reinforcement learning baselines across three multimodal LLM backbones. Notably, on Qwen3-VL-8B-Instruct, it surpasses HuatuoGPT-Vision-34B by 2.79 points and reduces early reasoning failure rates from 64.0% to 13.0%.
This study addresses the lack of effective evaluation for breadth-first search capabilities—specifically, exhaustive enumeration of members within a closed set and their structured attributes—in non-English contexts. It presents the first Korean-language benchmark for breadth-oriented agent assessment, leveraging an automated synthesis and validation pipeline to generate tasks that require agents to fully enumerate members of a given parent entity and populate their structured attribute tables. The work introduces a novel structured difficulty modulation mechanism, controlling table width and two-dimensional composite keys, alongside a unified normalized matcher and a multidimensional scoring framework (Item-, Column-, and Row-F1). Experiments reveal that while current agents achieve strong member identification (Item-F1: 92.8), they struggle significantly with complete row completion (Row-F1: 53.7), with performance degrading as task difficulty increases, thereby underscoring the benchmark’s necessity and challenge.
This work addresses the critical issue that existing biomedical agents often generate responses inconsistent with cited literature, a flaw overlooked by conventional evaluations reliant on fixed-answer benchmarks. To tackle this, the authors introduce the first retrieval-based agent evaluation benchmark for open-ended biomedical questions, comprising 12,553 unresolved queries spanning 12 domains. Agents must perform multi-turn tool-augmented reasoning and autonomously decide when to abstain from answering, enabling assessment of both factual faithfulness and refusal capability. Key innovations include an open-question paradigm without predefined answers, validation of question openness via subsequent real-world literature, objective difficulty calibration based on reference model failure rates, and a frozen checklist that significantly improves annotation consistency (Spearman correlation rising from 0.35 to 0.82). Experiments reveal that even state-of-the-art agents achieve only 29%–60% success on the hardest subset and exhibit pronounced tool-use degradation under high difficulty.
本文提出了一种计算高效的两步超参数转移框架,通过跨模型宽度和令牌维度转移学习率,解决大规模混合专家模型的优化问题。
This work addresses the challenge in reinforcement learning for instruction following where static prompts mismatch the policy’s evolving capabilities, resulting in poorly discriminative reward signals. To resolve this, the authors propose the LLM-as-a-Tutor framework, which treats prompt adaptation as a policy-aware process. Leveraging a large language model as both examiner and tutor, the method identifies non-challenging prompts through pairwise comparisons and monotonically increases task difficulty by appending atomic constraints—ensuring training signals remain calibrated to the policy’s current proficiency. Notably, this approach requires no external scheduling mechanism and significantly outperforms policy-agnostic baselines and existing adaptive methods across three complex instruction-following benchmarks.
This work addresses a critical limitation in existing post-training methods for medical multimodal large language models, which overly prioritize final answer correctness while neglecting the optimization of intermediate reasoning steps, thereby suffering from cascading failures triggered by early errors. To mitigate this, we propose Medical Reasoning-aware Policy Optimization (MRPO), the first reinforcement learning framework tailored for medical multimodal reasoning that incorporates step-aware feedback. MRPO introduces step-level process rewards, applying exponentially stronger penalties to ineffective early reasoning even when the final answer is incorrect, thus effectively interrupting error propagation. Experimental results demonstrate that MRPO consistently outperforms standard GRPO and state-of-the-art reinforcement learning baselines across three multimodal LLM backbones. Notably, on Qwen3-VL-8B-Instruct, it surpasses HuatuoGPT-Vision-34B by 2.79 points and reduces early reasoning failure rates from 64.0% to 13.0%.
This study addresses the lack of effective evaluation for breadth-first search capabilities—specifically, exhaustive enumeration of members within a closed set and their structured attributes—in non-English contexts. It presents the first Korean-language benchmark for breadth-oriented agent assessment, leveraging an automated synthesis and validation pipeline to generate tasks that require agents to fully enumerate members of a given parent entity and populate their structured attribute tables. The work introduces a novel structured difficulty modulation mechanism, controlling table width and two-dimensional composite keys, alongside a unified normalized matcher and a multidimensional scoring framework (Item-, Column-, and Row-F1). Experiments reveal that while current agents achieve strong member identification (Item-F1: 92.8), they struggle significantly with complete row completion (Row-F1: 53.7), with performance degrading as task difficulty increases, thereby underscoring the benchmark’s necessity and challenge.
This work addresses the critical issue that existing biomedical agents often generate responses inconsistent with cited literature, a flaw overlooked by conventional evaluations reliant on fixed-answer benchmarks. To tackle this, the authors introduce the first retrieval-based agent evaluation benchmark for open-ended biomedical questions, comprising 12,553 unresolved queries spanning 12 domains. Agents must perform multi-turn tool-augmented reasoning and autonomously decide when to abstain from answering, enabling assessment of both factual faithfulness and refusal capability. Key innovations include an open-question paradigm without predefined answers, validation of question openness via subsequent real-world literature, objective difficulty calibration based on reference model failure rates, and a frozen checklist that significantly improves annotation consistency (Spearman correlation rising from 0.35 to 0.82). Experiments reveal that even state-of-the-art agents achieve only 29%–60% success on the hardest subset and exhibit pronounced tool-use degradation under high difficulty.