PyraMathBench: Evaluating and Improving Mathematical Capability in Large Language Models

Zetian Ouyang; Linlin Wang; Gerard De Melo; Liang He

PyraMathBench: Evaluating and Improving Mathematical Capability in Large Language Models

Zetian Ouyang, Linlin Wang, Gerard de Melo, Liang He

Abstract

Despite the pivotal role of numerical reasoning as the cornerstone of mathematical capabilities in large language models (LLMs) across applications, few benchmarks evaluate LLMs by integrating numerical processing and mathematical reasoning, hindering the interpretability of failures in math tasks. We introduce PyraMathBench, a comprehensive hierarchical benchmark with 27,215 questions derived from 7,404 math word problems, spanning 4 key cognitive aspects, 14 subcategories, and 2 modalities. Experiments reveal that LLMs’ performance is severely compromised by inadequate numerical computation and weak handling of abstract numerical questions. To address this, we propose the Smart Optimization Learning-based VErsatile module (SOLVE) and Interactive Relative Policy Optimization (IRPO), which enhance LLMs’ numerical-mathematical synergy via efficient tool calls (fuzzy matching and low-quality call rejection). Comparative experiments show Qwen-2.5 achieves a 5.0 score improvement with SOLVE and IRPO training.

Anthology ID:: 2026.findings-acl.1869
Volume:: Findings of the Association for Computational Linguistics: ACL 2026
Month:: July
Year:: 2026
Address:: San Diego, California, United States
Editors:: Maria Liakata, Viviane P. Moreira, Jiajun Zhang, David Jurgens
Venue:: Findings
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 37481–37511
Language:
URL:: https://preview.aclanthology.org/ingest-acl/2026.findings-acl.1869/
DOI:
Bibkey:
Cite (ACL):: Zetian Ouyang, Linlin Wang, Gerard de Melo, and Liang He. 2026. PyraMathBench: Evaluating and Improving Mathematical Capability in Large Language Models. In Findings of the Association for Computational Linguistics: ACL 2026, pages 37481–37511, San Diego, California, United States. Association for Computational Linguistics.
Cite (Informal):: PyraMathBench: Evaluating and Improving Mathematical Capability in Large Language Models (Ouyang et al., Findings 2026)
Copy Citation:
PDF:: https://preview.aclanthology.org/ingest-acl/2026.findings-acl.1869.pdf
Checklist:: 2026.findings-acl.1869.checklist.pdf

PDF Cite Search Checklist Fix data