arXiv:cs.AI(全量分类)· Khawaja Murad ul Hassan, Mehran Ebrahimi·· 5 小时前AI 评分38
推理时 PRM 剪枝片段嫁接为何失效:来自三个推理 LM 的证据
Characterizing a Configuration Where Inference-Time PRM-Pruned Fragment Grafting Is Inert: Evidence from Three Reasoning LMs
AI 导读
研究将 PRM-Pruned Fragment Grafting(PPFG)这一推理时干预机制单独隔离测试,在 Qwen2.5-7B-Instruct 搭配 Math-Shepherd、完整 MATH500(n=500,三个随机种子)上,其停滞与随机目标两种变体在所有测量维度均与独立并行 CoT 基线统计上无法区分。
正文
Abstract:Diversity collapse in parallel chain-of-thought has motivated inference-time interventions built on a natural design: when a process reward model (PRM) prunes a chain, its high-PRM prefix is extracted and grafted verbatim as an in-context demonstration into a still-decoding sibling. We isolate this mechanism, PRM-Pruned Fragment Grafting (PPFG), as the most cost-minimal operationalization of cross-trajectory step-level transfer, and test it at the operating point where prior fragment-grafting work reports gains only under additional compensating ingredients. On Qwen2.5-7B-Instruct with Math-Shepherd on full MATH500 (n=500, three seeds), PPFG in both stagnation- and random-targeting variants is statistically indistinguishable from an independent parallel-CoT baseline on every measured axis. We characterize why: a four-bucket classification of 322 stagnation-rule injection events shows only 14% targeted a genuinely struggling chain; the rest landed on chains that had already succeeded, were near completion, or sat on a flat PRM plateau, states a rescue graft cannot change. No compound-gate refinement jointly achieves well-targeted firing and adequate density, and a random control matches the same parity at 2.4x the firing rate, so the inertness is not heuristic-specific. The finding replicates across three base LMs, six benchmarks, a second PRM, and a compatibility-gate sweep; two-one-sided-tests analysis promotes the parity to positive equivalence on all twelve Qwen/LLaMA cells. A per-event spot-check finds injected chains prune at 2.75x the matched-step rate, but a surviving-sibling counterfactual finds no population-level compensation. A hindsight oracle bounds any per-problem gain from choosing PPFG over independent at +0.13 pp. We contribute an equivalence-testing template for establishing inference-time mechanism nulls, with every claim scoped to its tested operating point.
| Comments: | 24 pages, 4 figures, 22 tables |
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.00047 [cs.AI] |
| (or arXiv:2610.00047v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00047 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Khawaja Murad Ul Hassan [view email]
[v1]
Thu, 3 Sep 2026 11:06:07 UTC (117 KB)
来源:arXiv:cs.AI(全量分类) · arxiv.org