跳到正文
arXiv:cs.CL· Yuren Hao, Xiang Wan, ChengXiang Zhai·· 3 小时前

LLM 数学推理鲁棒性研究:用数学等价变换对高等数学题进行基准测试

An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems

AI 导读

研究者提出 GAP(Generalisation-and-Perturbation)方法,通过表面重命名和内核重写两类可解释变换自动生成数学等价变体;基于 1938 至 2024 年全部 1,051 道 Putnam 竞赛题构建了 6,306 题的 PutnamGAP 语料,含 5,255 个未见变体。

正文

View PDF HTML (experimental)

Abstract:Frontier large language models (LLMs) now reach near-ceiling accuracy on standard mathematical-reasoning benchmarks and gold-medal-level performance at the International Mathematical Olympiad. As these benchmarks saturate and their items leak into training data, a high score no longer shows whether a model reasons robustly or which component of its reasoning fails. To evaluate reasoning while keeping results informative and failures diagnosable, we propose GAP (Generalisation-and-Perturbation), a methodology that automatically generates mathematically equivalent variants of existing mathematics problems at scale using two disjoint, interpretable transformations: (1) surface renames, probing the binding between identifiers and latent variable roles, and (2) kernel rewrites, probing whether a high-level proof plan survives a change of mathematical setting. Compared with existing benchmarks, GAP has two key benefits: (1) novel, likely unseen variants mitigate data leakage, and (2) performance across transformation families enables failure diagnosis, each transformation testing a hypothesis about the cause of failure. We instantiate GAP on all 1,051 William Lowell Putnam Competition problems from 1938 to 2024, adding 5,255 unseen variants to form PutnamGAP, a 6,306-item competition-level mathematics corpus and the first public machine-readable dataset from the full Putnam archive. Using PutnamGAP, we evaluated 18 commercial and open-source models spanning sizes and providers. Accuracy drops across all models and variant families, most severely under kernel rewrites. This gap does not close with model strength, suggesting that even the strongest models' dominant weakness is transferring a proof plan to a changed mathematical setting, rather than handling surface changes. Further analysis provides finer failure diagnoses and potentially useful insights for improving LLM reasoning.
Comments: 34 pages, 9 figures, accepted at ICLR 2026 workshop
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
MSC classes: 68T50
ACM classes: I.2.7
Cite as: arXiv:2508.08833 [cs.CL]
  (or arXiv:2508.08833v4 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2508.08833

arXiv-issued DOI via DataCite

Submission history

From: Yuren Hao [view email]
[v1] Tue, 12 Aug 2025 10:40:33 UTC (2,097 KB)
[v2] Wed, 8 Oct 2025 01:46:12 UTC (3,354 KB)
[v3] Thu, 4 Dec 2025 08:10:06 UTC (3,368 KB)
[v4] Wed, 7 Oct 2026 22:06:35 UTC (4,846 KB)

来源:arXiv:cs.CL · arxiv.org