MameLoshnLM:意第绪语语言模型与评估基准
MameLoshnLM: Yiddish Language Model and Evaluation Benchmark
August 6, 2026
作者: Uri Katz, Omer Goldman, Tomasz Limisiewicz, Reut Tsarfaty, Noah A. Smith
cs.AI
摘要
我们提出了 MameLoshnLM,这是首个专为意第绪语构建的开源 8B 参数语言模型。尽管意第绪语拥有丰富的文本传统,但其有限的数字存在以及可靠评估资源的匮乏,制约了意第绪语语言建模的进展。现有的多语言语料库和基准往往不能很好地代表该语言,其中包含大量带有噪声、机器翻译和错误分类的文本。我们通过引入 Oytser(一个结合当代网络原生来源与文学材料的高质量意第绪语预训练语料库)和 Kashes(一个涵盖翻译、语言分析、信息抽取和语言理解的多任务基准)来弥补这些不足。利用这些资源,我们对 Llama 3.1 8B 进行继续预训练,得到 MameLoshnLM。在该基准的各项任务中,MameLoshnLM 均优于同等规模的开源基线模型。我们的分析表明,这些提升不仅是数量上的:与通用多语言模型相比,MameLoshnLM 能更好地捕捉定义了该语言特征的词汇和形态模式,从而揭示了嘈杂的网络规模多语言数据在低资源语言上更普遍的失败模式。我们的研究结果既为意第绪语自然语言处理奠定了基础,也为在历史丰富但数字代表性不足的语言中开发语言模型提供了实用模板。
English
We present MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish. Despite Yiddish's rich textual tradition, its limited digital presence and the scarcity of reliable evaluation resources have constrained progress in Yiddish language modeling. Existing multilingual corpora and benchmarks are often poor proxies for the language, containing substantial amounts of noisy, machine-translated, and misclassified text. We address these gaps by introducing Oytser, a high-quality Yiddish pretraining corpus that combines contemporary web-native sources with literary materials, and Kashes, a multi-task benchmark spanning translation, linguistic analysis, information extraction, and language understanding. Using these resources, we continue pretraining Llama 3.1 8B to obtain MameLoshnLM. Across the tasks in the benchmark, MameLoshnLM outperforms open baselines of similar scale. Our analyses show that these gains are not only quantitative: relative to general-purpose multilingual models, MameLoshnLM better captures language-defining lexical and morphological patterns, pointing to a broader failure mode of noisy web-scale multilingual data for low-resource languages. Our results provide both a foundation for Yiddish NLP and a practical template for language model development in historically rich but digitally underrepresented languages.