MameLoshnLM: 이디시어 언어 모델 및 평가 벤치마크
MameLoshnLM: Yiddish Language Model and Evaluation Benchmark
August 6, 2026
저자: Uri Katz, Omer Goldman, Tomasz Limisiewicz, Reut Tsarfaty, Noah A. Smith
cs.AI
초록
우리는 이디시어를 위해 특별히 구축된 최초의 오픈소스 80억 파라미터 언어 모델인 MameLoshnLM을 제시한다. 이디시어의 풍부한 텍스트 전통에도 불구하고, 제한된 디지털 존재감과 신뢰할 수 있는 평가 자원의 부족은 이디시어 언어 모델링의 발전을 제약해 왔다. 기존의 다국어 코퍼스와 벤치마크는 종종 해당 언어에 대한 부실한 대리물로서, 상당량의 노이즈가 섞인 기계 번역 텍스트와 잘못 분류된 텍스트를 포함하고 있다. 우리는 이러한 공백을 해결하기 위해 현대의 웹 기반 자료와 문학 자료를 결합한 고품질 이디시어 사전학습 코퍼스인 Oytser와 번역, 언어 분석, 정보 추출, 언어 이해를 아우르는 다중 작업 벤치마크인 Kashes를 소개한다. 이러한 자원을 사용하여 우리는 Llama 3.1 8B에 대한 사전학습을 지속함으로써 MameLoshnLM을 얻는다. 벤치마크의 작업 전반에 걸쳐 MameLoshnLM은 유사한 규모의 오픈 베이스라인을 능가한다. 우리의 분석은 이러한 성과가 단지 양적인 것만이 아님을 보여준다: 범용 다국어 모델과 비교할 때, MameLoshnLM은 언어를 정의하는 어휘적 및 형태적 패턴을 더 잘 포착하며, 이는 저자원 언어에 대한 노이즈가 있는 웹 규모 다국어 데이터의 광범위한 실패 양상을 시사한다. 우리의 결과는 이디시어 NLP를 위한 기반과 역사적으로 풍부하지만 디지털에서 과소 대표되는 언어들에서의 언어 모델 개발을 위한 실용적인 템플릿을 모두 제공한다.
English
We present MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish. Despite Yiddish's rich textual tradition, its limited digital presence and the scarcity of reliable evaluation resources have constrained progress in Yiddish language modeling. Existing multilingual corpora and benchmarks are often poor proxies for the language, containing substantial amounts of noisy, machine-translated, and misclassified text. We address these gaps by introducing Oytser, a high-quality Yiddish pretraining corpus that combines contemporary web-native sources with literary materials, and Kashes, a multi-task benchmark spanning translation, linguistic analysis, information extraction, and language understanding. Using these resources, we continue pretraining Llama 3.1 8B to obtain MameLoshnLM. Across the tasks in the benchmark, MameLoshnLM outperforms open baselines of similar scale. Our analyses show that these gains are not only quantitative: relative to general-purpose multilingual models, MameLoshnLM better captures language-defining lexical and morphological patterns, pointing to a broader failure mode of noisy web-scale multilingual data for low-resource languages. Our results provide both a foundation for Yiddish NLP and a practical template for language model development in historically rich but digitally underrepresented languages.