跳到正文
原文
arXiv:cs.AI(全量分类)· Khawaja Murad ul Hassan, Mehran Ebrahimi·· 5 小时前AI 评分38

推理时 PRM 剪枝片段嫁接为何失效:来自三个推理 LM 的证据

Characterizing a Configuration Where Inference-Time PRM-Pruned Fragment Grafting Is Inert: Evidence from Three Reasoning LMs

AI 导读

研究将 PRM-Pruned Fragment Grafting(PPFG)这一推理时干预机制单独隔离测试,在 Qwen2.5-7B-Instruct 搭配 Math-Shepherd、完整 MATH500(n=500,三个随机种子)上,其停滞与随机目标两种变体在所有测量维度均与独立并行 CoT 基线统计上无法区分。

正文

View PDF HTML (experimental)

Abstract:Diversity collapse in parallel chain-of-thought has motivated inference-time interventions built on a natural design: when a process reward model (PRM) prunes a chain, its high-PRM prefix is extracted and grafted verbatim as an in-context demonstration into a still-decoding sibling. We isolate this mechanism, PRM-Pruned Fragment Grafting (PPFG), as the most cost-minimal operationalization of cross-trajectory step-level transfer, and test it at the operating point where prior fragment-grafting work reports gains only under additional compensating ingredients. On Qwen2.5-7B-Instruct with Math-Shepherd on full MATH500 (n=500, three seeds), PPFG in both stagnation- and random-targeting variants is statistically indistinguishable from an independent parallel-CoT baseline on every measured axis. We characterize why: a four-bucket classification of 322 stagnation-rule injection events shows only 14% targeted a genuinely struggling chain; the rest landed on chains that had already succeeded, were near completion, or sat on a flat PRM plateau, states a rescue graft cannot change. No compound-gate refinement jointly achieves well-targeted firing and adequate density, and a random control matches the same parity at 2.4x the firing rate, so the inertness is not heuristic-specific. The finding replicates across three base LMs, six benchmarks, a second PRM, and a compatibility-gate sweep; two-one-sided-tests analysis promotes the parity to positive equivalence on all twelve Qwen/LLaMA cells. A per-event spot-check finds injected chains prune at 2.75x the matched-step rate, but a surviving-sibling counterfactual finds no population-level compensation. A hindsight oracle bounds any per-problem gain from choosing PPFG over independent at +0.13 pp. We contribute an equivalence-testing template for establishing inference-time mechanism nulls, with every claim scoped to its tested operating point.
Comments: 24 pages, 4 figures, 22 tables
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as: arXiv:2610.00047 [cs.AI]
  (or arXiv:2610.00047v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.00047

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Khawaja Murad Ul Hassan [view email]
[v1] Thu, 3 Sep 2026 11:06:07 UTC (117 KB)

来源:arXiv:cs.AI(全量分类) · arxiv.org