arXiv:cs.CL· Yuren Hao, Xiang Wan, ChengXiang Zhai·· 3 小时前
LLM 数学推理鲁棒性研究:用数学等价变换对高等数学题进行基准测试
An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
AI 导读
研究者提出 GAP(Generalisation-and-Perturbation)方法,通过表面重命名和内核重写两类可解释变换自动生成数学等价变体;基于 1938 至 2024 年全部 1,051 道 Putnam 竞赛题构建了 6,306 题的 PutnamGAP 语料,含 5,255 个未见变体。
正文
Abstract:Frontier large language models (LLMs) now reach near-ceiling accuracy on standard mathematical-reasoning benchmarks and gold-medal-level performance at the International Mathematical Olympiad. As these benchmarks saturate and their items leak into training data, a high score no longer shows whether a model reasons robustly or which component of its reasoning fails. To evaluate reasoning while keeping results informative and failures diagnosable, we propose GAP (Generalisation-and-Perturbation), a methodology that automatically generates mathematically equivalent variants of existing mathematics problems at scale using two disjoint, interpretable transformations: (1) surface renames, probing the binding between identifiers and latent variable roles, and (2) kernel rewrites, probing whether a high-level proof plan survives a change of mathematical setting. Compared with existing benchmarks, GAP has two key benefits: (1) novel, likely unseen variants mitigate data leakage, and (2) performance across transformation families enables failure diagnosis, each transformation testing a hypothesis about the cause of failure. We instantiate GAP on all 1,051 William Lowell Putnam Competition problems from 1938 to 2024, adding 5,255 unseen variants to form PutnamGAP, a 6,306-item competition-level mathematics corpus and the first public machine-readable dataset from the full Putnam archive. Using PutnamGAP, we evaluated 18 commercial and open-source models spanning sizes and providers. Accuracy drops across all models and variant families, most severely under kernel rewrites. This gap does not close with model strength, suggesting that even the strongest models' dominant weakness is transferring a proof plan to a changed mathematical setting, rather than handling surface changes. Further analysis provides finer failure diagnoses and potentially useful insights for improving LLM reasoning.
| Comments: | 34 pages, 9 figures, accepted at ICLR 2026 workshop |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| MSC classes: | 68T50 |
| ACM classes: | I.2.7 |
| Cite as: | arXiv:2508.08833 [cs.CL] |
| (or arXiv:2508.08833v4 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2508.08833 arXiv-issued DOI via DataCite |
Submission history
From: Yuren Hao [view email]
[v1]
Tue, 12 Aug 2025 10:40:33 UTC (2,097 KB)
[v2]
Wed, 8 Oct 2025 01:46:12 UTC (3,354 KB)
[v3]
Thu, 4 Dec 2025 08:10:06 UTC (3,368 KB)
[v4]
Wed, 7 Oct 2026 22:06:35 UTC (4,846 KB)
来源:arXiv:cs.CL · arxiv.org