PluraMath: 고자원 언어 너머의 수학적 추론 평가 확장
PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages
July 7, 2026
저자: Daryna Dementieva, Nikolay Babakov, Kathy Hämmerl, Ilseyar Alimova, Jindřich Libovický, Shu Okabe, Miras Baisbay, Lukas Edman, Abrorkhon Inomkhujaev, Antonia Karamolegkou, Mateusz Lango, Volkan Özer, Nikola Selic, Subhankar Swain, Tsedeniya Kinfe Temesgen, Galit Bary Weisberg, Alexander Fraser
cs.AI
초록
수학적 추론은 추론용 대규모 언어 모델(LLMs)을 평가하고 조정하기 위한 핵심 과제가 되었지만, 기존 벤치마크는 사전 학습 코퍼스와 평가 스위트 모두에서 영어와 중국어가 지배하는 고자원 언어에 크게 편향되어 있다. 최근 공개된 PolyMath(Wang et al., 2025) 데이터셋은 중요한 진전을 나타내지만, 그 범위는 여전히 18개의 고자원 언어로 제한되어 있다. 이러한 격차를 해소하기 위해 우리는 PolyMath를 6개 어족에 속하는 18개의 추가 저대표 언어(중간 자원부터 극저자원 설정까지 포함)로 확장한 PluraMath를 소개한다. 우리는 원어민이 사전 계산된 번역을 철저히 검증하는 인간 큐레이션 파이프라인을 통해 데이터셋을 구축했다. 이후 PluraMath를 사용하여 소형, 중형, 대형, 폐쇄형 앙상블의 네 가지 모델 규모에 걸쳐 27개의 추론용 LLM을 벤치마킹함으로써, 다양한 언어 조건에서 최신 모델의 다국어 수학적 추론 능력을 탐구했다. 우리의 세부 분석은 고자원 언어와 저대표 언어 간 수학적 추론 성능에 지속적인 격차가 존재함을 확인했으며, 더 강력한 결과는 대체로 더 나은 명령 수행 능력과 관련이 있었다. 우리는 데이터셋, 데이터 수집 파이프라인, 평가 프레임워크를 완전히 오픈소스로 공개하여, 저대표 커뮤니티의 다국어 벤치마크 개발 장벽을 낮추는 것을 목표로 한다.
English
Mathematical reasoning has become a central task for evaluating and tuning reasoning Large Language Models (LLMs), yet existing benchmarks remain heavily biased toward high-resource languages, with English and Chinese dominating both pre-training corpora and evaluation suites. The recently released PolyMath (Wang et al., 2025) dataset represents a significant step forward, yet its coverage is still limited to 18 only high-resource languages. To address this gap, we introduce PluraMath, an extension of PolyMath to 18 additional {underrepresented languages spanning 6 language families -- ranging from mid-resource to extreme low-resource settings. We constructed the dataset through a human-curated pipeline, where native speakers thoroughly validated pre-computed translations. Using PluraMath, we then benchmark 27 reasoning LLMs across four model scales -- small, mid-size, large, and closed-source ensembles -- probing the multilingual mathematical reasoning capabilities of state-of-the-art models under diverse linguistic conditions. Our fine-grained analysis confirms a persistent gap in mathematical reasoning performance between high-resource and underrepresented languages, with stronger results largely associated with better instruction-following ability. We fully open-source our dataset, data acquisition pipeline, and evaluation framework, with the goal of lowering the barrier to multilingual benchmark development for underrepresented communities.