DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data

📅 2026-08-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the prevalent reliance of large language models on non-compliant data, which hinders research aligned with open-source and ethical standards. The authors propose a 1-billion-parameter language model trained entirely from scratch using only 161 compliant datasets within a Hierarchical Reasoning Model (HRM) architecture, covering English, mathematics, code, and Danish. The resulting model outperforms the original HRM-Text 1B across all 20 benchmark evaluations and establishes a new state-of-the-art for 1B-scale models on Danish-language tasks. Its performance rivals that of significantly larger models such as Qwen 3.5 4B and Gemma 4 E2B. The model has been openly released on Hugging Face to support transparent and responsible research.
📝 Abstract
Current large language model development relies on massive, often non-permissible datasets, creating a high barrier for researchers committed to open-source and ethically sourced data. We introduce Mimir v1, a 1-billion-parameter language model based on the Hierarchical Reasoning Model (HRM) architecture, that is trained from scratch and delivers highly competitive performance for English and sets a new state of the art for Danish using only permissible post-training data. Trained on a mixture of 161 datasets, Mimir v1 outperforms the original HRM-Text 1B and competes with larger frontier models like Qwen 3.5 4B and Gemma 4 E2B, tested across 20 benchmarks for English, Math & Code and Danish. The model is available on the Hugging Face Hub: https://huggingface.co/danish-foundation-models/DFM-Mimir
Problem

Research questions and friction points this paper is trying to address.

permissible data
large language models
ethical data sourcing
open-source AI
data compliance
Innovation

Methods, ideas, or system contributions that make the work stand out.

permissible data
Hierarchical Reasoning Model
open-source LLM
ethical AI
multilingual benchmarking
🔎 Similar Papers