๐ค AI Summary
To address the challenge of balancing complex multi-step reasoning with open-ended instruction following in large language models, this paper introduces the Hermes-4 series of hybrid reasoning models. Methodologically, we propose a unified training framework that jointly optimizes structured multi-turn reasoning and open instruction comprehension, leveraging large-scale cleaned corpora and high-quality synthetic dataโincluding behavior-guided reasoning trajectories. A multi-dimensional evaluation suite ensures training alignment and fidelity. Our key contribution is the first end-to-end unified modeling of structured reasoning chains (e.g., mathematical derivation, program debugging) and general-purpose instruction understanding. Hermes-4 achieves significant improvements over same-scale baselines on benchmarks including GSM8K, HumanEval, and MMLU. We fully open-source the model weights and detailed training configurations to foster reproducible research in hybrid reasoning.
๐ Abstract
We present Hermes 4, a family of hybrid reasoning models that combine structured, multi-turn reasoning with broad instruction-following ability. We describe the challenges encountered during data curation, synthesis, training, and evaluation, and outline the solutions employed to address these challenges at scale. We comprehensively evaluate across mathematical reasoning, coding, knowledge, comprehension, and alignment benchmarks, and we report both quantitative performance and qualitative behavioral analysis. To support open research, all model weights are published publicly at https://huggingface.co/collections/NousResearch/hermes-4-collection-68a731bfd452e20816725728