PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages
MCML Authors
Abstract
Abstract
Mathematical reasoning has become a central task for evaluating and tuning reasoning Large Language Models (LLMs), yet existing benchmarks remain heavily biased toward high-resource languages, with English and Chinese dominating both pre-training corpora and evaluation suites. The recently released PolyMath (Wang et al., 2025) dataset represents a significant step forward, yet its coverage is still limited to 18 only high-resource languages. To address this gap, we introduce PLURAMATH, an extension of PolyMath to 18 additional underrepresented languages spanning 6 language families—ranging from mid-resource to extreme low-resource settings. We constructed the dataset through a human-curated pipeline, where native speakers thoroughly validated precomputed translations. Using PLURAMATH, we then benchmark 27 reasoning LLMs across<br>four model scales—small, mid-size, large, and closed-source ensembles—probing the multilingual mathematical reasoning capabilities of state-of-the-art models under diverse linguistic conditions. Our fine-grained analysis confirms a persistent gap in mathematical reasoning performance between high-resource and underrepresented languages, with stronger results largely associated with better instructionfollowing ability. We fully open-source our dataset, data acquisition pipeline, and evaluation framework, with the goal of lowering the barrier to multilingual benchmark development for underrepresented communities.
misc DBH+26
Preprint
Jul. 2026Authors
D. Dementieva • N. Babakov • K. Hämmerl • I. Alimova • J. Libovický • S. Okabe • M. Baisbay • L. Edman • A. Inomkhujaev • A. Karamolegkou • M. Lango • V. Özer • N. Selic • S. Swain • T. K. Temesgen • G. B. Weisberg • A. FraserLinks
arXiv GitHubResearch Area
BibTeXKey: DBH+26