MameLoshnLM:意第緒語語言模型與評估基準
MameLoshnLM: Yiddish Language Model and Evaluation Benchmark
August 6, 2026
作者: Uri Katz, Omer Goldman, Tomasz Limisiewicz, Reut Tsarfaty, Noah A. Smith
cs.AI
摘要
我們提出了MameLoshnLM,這是第一個專為意第緒語建構的開源80億參數語言模型。儘管意第緒語擁有豐富的文本傳統,但其有限的數位足跡與可靠評測資源的匱乏,制約了意第緒語語言建模的進展。現有的多語言語料庫與基準評測往往難以忠實反映該語言,其中包含大量雜訊文本、機器翻譯文本及誤分類文本。為填補這些缺口,我們引入了Oytser——一個高品質的意第緒語預訓練語料庫,結合了當代網路原生來源與文學材料——以及Kashes——一個涵蓋翻譯、語言學分析、資訊抽取與語言理解的多任務基準評測。利用這些資源,我們對Llama 3.1 8B進行持續預訓練,從而獲得MameLoshnLM。在基準評測的各項任務中,MameLoshnLM均優於同量級的开源基線模型。我們的分析表明,這些提升不僅體現在量化指標上:相較於通用型多語言模型,MameLoshnLM能更準確地捕捉定義語言特徵的詞彙與形態模式,這指向了雜訊網路規模多語言資料在低資源語言上的更廣泛失效模式。我們的研究成果既為意第緒語自然語言處理奠定了基礎,也為歷史上底蘊深厚但數位資源不足的語言提供了語言模型開發的實用範本。
English
We present MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish. Despite Yiddish's rich textual tradition, its limited digital presence and the scarcity of reliable evaluation resources have constrained progress in Yiddish language modeling. Existing multilingual corpora and benchmarks are often poor proxies for the language, containing substantial amounts of noisy, machine-translated, and misclassified text. We address these gaps by introducing Oytser, a high-quality Yiddish pretraining corpus that combines contemporary web-native sources with literary materials, and Kashes, a multi-task benchmark spanning translation, linguistic analysis, information extraction, and language understanding. Using these resources, we continue pretraining Llama 3.1 8B to obtain MameLoshnLM. Across the tasks in the benchmark, MameLoshnLM outperforms open baselines of similar scale. Our analyses show that these gains are not only quantitative: relative to general-purpose multilingual models, MameLoshnLM better captures language-defining lexical and morphological patterns, pointing to a broader failure mode of noisy web-scale multilingual data for low-resource languages. Our results provide both a foundation for Yiddish NLP and a practical template for language model development in historically rich but digitally underrepresented languages.