跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Shuo Xing, Zilin Dai, Chengyuan Qian, Fangzhou Lin, Wenjing Chen, Ping He, Pan Lu, Alvaro Velasquez, Mohit Bansal, Zhengzhong Tu·· 14 小时前AI 评分37

缺失的原语:诊断并修复大语言模型的数学推理能力

The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models

AI 导读

研究者提出"数学原语"概念及 \hlei{} 基准,从 Discovery、Generation、Digestion、Execution 四个维度评估 LLM 的数学推理,发现解题准确率掩盖了不同的能力画像,Discovery 是数学推理的主要瓶颈。基于此提出 \abs{} 原语优先自蒸馏框架,将原语引导的推理选择性迁移至学生模型,在多种模型规模和挑战性基准上持续提升数学推理表现。

正文

View PDF HTML (experimental)

Abstract:While Large Language Models (LLMs) have demonstrated striking capabilities on frontier mathematical problems, it remains unclear whether they possess the structural mathematical understanding underlying their solutions. In this paper, we take a first step toward systematically studying mathematical understanding in LLMs, from diagnosing its distinct capabilities to leveraging these findings to improve post-training. First, we introduce the notion of Mathematical Primitive to probe structural mathematical understanding and propose \hlei{}, a novel benchmark that evaluates mathematical reasoning along four distinct dimensions: Discovery, Generation, Digestion, and Execution. Second, our systematic diagnosis shows that solution accuracy masks distinct capability profiles, primitives unlock substantial latent execution capacity, and Discovery is the dominant bottleneck in mathematical reasoning. Our post-training analysis further shows that discovery-limited failures are particularly amenable to repair. Finally, building on these findings, we introduce \abs{}, a primitive-privileged self-distillation framework that selectively transfers primitive-guided reasoning into the student model. Extensive experiments demonstrate that \abs{} consistently improves mathematical reasoning over baselines across model scales and challenging benchmarks.
Comments: 27 pages
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.02191 [cs.LG]
  (or arXiv:2610.02191v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.02191

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Shuo Xing [view email]
[v1] Thu, 1 Oct 2026 17:59:32 UTC (2,696 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org