PluraMath: 高リソース言語を超えた数学的推論評価の拡張
PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages
July 7, 2026
著者: Daryna Dementieva, Nikolay Babakov, Kathy Hämmerl, Ilseyar Alimova, Jindřich Libovický, Shu Okabe, Miras Baisbay, Lukas Edman, Abrorkhon Inomkhujaev, Antonia Karamolegkou, Mateusz Lango, Volkan Özer, Nikola Selic, Subhankar Swain, Tsedeniya Kinfe Temesgen, Galit Bary Weisberg, Alexander Fraser
cs.AI
要旨
数学的推論は、推論を行う大規模言語モデル(LLM)の評価とチューニングにおいて中心的なタスクとなっている。しかし、既存のベンチマークは依然として高リソース言語に大きく偏っており、英語と中国語が事前学習コーパスと評価スイートの両方で支配的である。最近公開されたPolyMathデータセット(Wang et al., 2025)は大きな前進ではあるが、そのカバレッジは依然として18の高リソース言語のみに限られている。このギャップに対処するため、我々はPluraMathを導入する。これはPolyMathを拡張し、18の追加の過小評価言語(6つの語族にわたり、中リソースから極低リソースまで)をカバーする。データセットは人間によるキュレーションパイプラインを通じて構築され、ネイティブスピーカーが事前計算された翻訳を徹底的に検証した。PluraMathを用いて、我々は27の推論LLMを4つのモデル規模(小規模、中規模、大規模、クローズドソースのアンサンブル)にわたってベンチマークし、多様な言語条件下での最先端モデルの多言語数学的推論能力を調査した。我々の詳細な分析により、高リソース言語と過小評価言語の間の数学的推論性能における永続的なギャップが確認され、より良い結果は主に指示追従能力の高さと関連していることが示された。我々は、データセット、データ取得パイプライン、評価フレームワークを完全にオープンソース化し、過小評価コミュニティにおける多言語ベンチマーク開発の障壁を下げることを目指している。
English
Mathematical reasoning has become a central task for evaluating and tuning reasoning Large Language Models (LLMs), yet existing benchmarks remain heavily biased toward high-resource languages, with English and Chinese dominating both pre-training corpora and evaluation suites. The recently released PolyMath (Wang et al., 2025) dataset represents a significant step forward, yet its coverage is still limited to 18 only high-resource languages. To address this gap, we introduce PluraMath, an extension of PolyMath to 18 additional {underrepresented languages spanning 6 language families -- ranging from mid-resource to extreme low-resource settings. We constructed the dataset through a human-curated pipeline, where native speakers thoroughly validated pre-computed translations. Using PluraMath, we then benchmark 27 reasoning LLMs across four model scales -- small, mid-size, large, and closed-source ensembles -- probing the multilingual mathematical reasoning capabilities of state-of-the-art models under diverse linguistic conditions. Our fine-grained analysis confirms a persistent gap in mathematical reasoning performance between high-resource and underrepresented languages, with stronger results largely associated with better instruction-following ability. We fully open-source our dataset, data acquisition pipeline, and evaluation framework, with the goal of lowering the barrier to multilingual benchmark development for underrepresented communities.