ChatPaper.aiChatPaper

MameLoshnLM:イディッシュ語言語モデルと評価ベンチマーク

MameLoshnLM: Yiddish Language Model and Evaluation Benchmark

August 6, 2026
著者: Uri Katz, Omer Goldman, Tomasz Limisiewicz, Reut Tsarfaty, Noah A. Smith
cs.AI

要旨

我々は、イディッシュ語に特化して構築された初のオープンソース80億パラメータ言語モデルであるMameLoshnLMを提案する。イディッシュ語は豊かな文書伝統を有するにもかかわらず、その限られたデジタル上の存在感と信頼性の高い評価リソースの不足により、イディッシュ語の言語モデリング研究の発展は制約されてきた。既存の多言語コーパスやベンチマークは、この言語にとって不十分な代替手段となることが多く、かなりの量のノイズを含む機械翻訳テキストや誤分類されたテキストが含まれている。我々はこれらのギャップに対処するため、現代のウェブ由来の情報源と文学作品の両方を組み合わせた高品質なイディッシュ語事前学習コーパスであるOytserと、翻訳、言語分析、情報抽出、言語理解にわたるマルチタスクベンチマークであるKashesを導入する。これらのリソースを用いて、Llama 3.1 8Bの継続事前学習を実施し、MameLoshnLMを獲得した。ベンチマーク内のタスク全体において、MameLoshnLMは同規模のオープンベースラインモデルを上回る性能を示す。我々の分析は、これらの改善が単に定量的なものではないことを示している。すなわち、汎用多言語モデルと比較して、MameLoshnLMは言語を特徴づける語彙的・形態的パターンをより適切に捉えており、これは低リソース言語に対するノイズの多いウェブ規模の多言語データのより広範な失敗モードを示唆している。本研究の成果は、イディッシュ語NLPの基盤を提供するとともに、歴史的に豊かでありながらデジタル上では過小評価されている言語に対する言語モデル開発の実践的な雛形を提供するものである。
English
We present MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish. Despite Yiddish's rich textual tradition, its limited digital presence and the scarcity of reliable evaluation resources have constrained progress in Yiddish language modeling. Existing multilingual corpora and benchmarks are often poor proxies for the language, containing substantial amounts of noisy, machine-translated, and misclassified text. We address these gaps by introducing Oytser, a high-quality Yiddish pretraining corpus that combines contemporary web-native sources with literary materials, and Kashes, a multi-task benchmark spanning translation, linguistic analysis, information extraction, and language understanding. Using these resources, we continue pretraining Llama 3.1 8B to obtain MameLoshnLM. Across the tasks in the benchmark, MameLoshnLM outperforms open baselines of similar scale. Our analyses show that these gains are not only quantitative: relative to general-purpose multilingual models, MameLoshnLM better captures language-defining lexical and morphological patterns, pointing to a broader failure mode of noisy web-scale multilingual data for low-resource languages. Our results provide both a foundation for Yiddish NLP and a practical template for language model development in historically rich but digitally underrepresented languages.