PluraMath:將數學推理評估擴展至高資源語言之外
PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages
July 7, 2026
作者: Daryna Dementieva, Nikolay Babakov, Kathy Hämmerl, Ilseyar Alimova, Jindřich Libovický, Shu Okabe, Miras Baisbay, Lukas Edman, Abrorkhon Inomkhujaev, Antonia Karamolegkou, Mateusz Lango, Volkan Özer, Nikola Selic, Subhankar Swain, Tsedeniya Kinfe Temesgen, Galit Bary Weisberg, Alexander Fraser
cs.AI
摘要
數學推理已成為評估和調校推理型大型語言模型(Reasoning LLMs)的核心任務,然而現有的基準測試仍嚴重偏向高資源語言,英語和中文在預訓練語料庫及評估套件中佔據主導地位。近期發表的 PolyMath 資料集(Wang 等人, 2025)雖是重要進展,但其涵蓋範圍仍僅限於 18 種高資源語言。為彌補此缺口,我們提出 PluraMath——將 PolyMath 擴展至另外 18 種涵蓋 6 個語系的低資源語言,範圍從中等資源到極低資源情境。我們透過人工策劃流程構建資料集,由母語人士徹底驗證預先計算的翻譯結果。利用 PluraMath,我們進一步對 27 個推理型 LLM 進行基準測試,涵蓋小型、中型、大型及封閉源整合模型等四種模型規模,探討最新模型在多樣語言條件下的多語言數學推理能力。我們的細粒度分析證實,高資源語言與低資源語言之間的數學推理表現持續存在差距,且較強的結果大致與較佳的指令遵循能力相關。我們將資料集、資料取得流程及評估框架完全開源,目標是降低低資源社群發展多語言基準測試的門檻。
English
Mathematical reasoning has become a central task for evaluating and tuning reasoning Large Language Models (LLMs), yet existing benchmarks remain heavily biased toward high-resource languages, with English and Chinese dominating both pre-training corpora and evaluation suites. The recently released PolyMath (Wang et al., 2025) dataset represents a significant step forward, yet its coverage is still limited to 18 only high-resource languages. To address this gap, we introduce PluraMath, an extension of PolyMath to 18 additional {underrepresented languages spanning 6 language families -- ranging from mid-resource to extreme low-resource settings. We constructed the dataset through a human-curated pipeline, where native speakers thoroughly validated pre-computed translations. Using PluraMath, we then benchmark 27 reasoning LLMs across four model scales -- small, mid-size, large, and closed-source ensembles -- probing the multilingual mathematical reasoning capabilities of state-of-the-art models under diverse linguistic conditions. Our fine-grained analysis confirms a persistent gap in mathematical reasoning performance between high-resource and underrepresented languages, with stronger results largely associated with better instruction-following ability. We fully open-source our dataset, data acquisition pipeline, and evaluation framework, with the goal of lowering the barrier to multilingual benchmark development for underrepresented communities.