ChatPaper.aiChatPaper

DFM Mimir v1:仅使用许可后训练数据、在10亿参数规模下实现前沿性能的开源HRM

DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data

August 13, 2026
作者: Peter Schneider-Kamp, Jacob Nielsen, Gianluca Barmina, Kenneth Enevoldsen, Lukas Galke Poech
cs.AI

摘要

当前大语言模型的开发依赖于大规模且往往未经许可的数据集,这为致力于开源和伦理合规数据来源的研究人员设置了较高的门槛。我们推出Mimir v1——一个基于层次推理模型(HRM)架构的10亿参数语言模型,该模型从零开始训练,在仅使用合规后训练数据的情况下,为英语带来极具竞争力的性能,并在丹麦语上刷新了最优水平。Mimir v1在161个数据集的混合数据上训练,性能优于原始的HRM-Text 1B,并能与更大的前沿模型(如Qwen 3.5 4B和Gemma 4 E2B)相抗衡,这一表现已在针对英语、数学与代码以及丹麦语的20项基准测试中得到验证。该模型可在Hugging Face Hub上获取:https://huggingface.co/danish-foundation-models/DFM-Mimir
English
Current large language model development relies on massive, often non-permissible datasets, creating a high barrier for researchers committed to open-source and ethically sourced data. We introduce Mimir v1, a 1-billion-parameter language model based on the Hierarchical Reasoning Model (HRM) architecture, that is trained from scratch and delivers highly competitive performance for English and sets a new state of the art for Danish using only permissible post-training data. Trained on a mixture of 161 datasets, Mimir v1 outperforms the original HRM-Text 1B and competes with larger frontier models like Qwen 3.5 4B and Gemma 4 E2B, tested across 20 benchmarks for English, Math & Code and Danish. The model is available on the Hugging Face Hub: https://huggingface.co/danish-foundation-models/DFM-Mimir