跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Saeed Ahmadnia, Cornelia Caragea·· 12 小时前AI 评分37

ReHoPER:用滚动时域规划增强大语言模型推理

ReHoPER: Receding-Horizon Planning for Enhanced Reasoning

AI 导读

ReHoPER 是一种仅推理、零样本的方法,通过在最终答案前沿多条路径生成并回答中间问题来提升大语言模型推理能力,它迭代规划一批候选中间问题、选择其一作答,再基于更新后的历史重新规划。

正文

View PDF HTML (experimental)

Abstract:We propose ReHoPER, an inference-only, zero-shot method that improves large language models' reasoning by generating and answering intermediate questions along multiple paths before the final answer. It iteratively plans a horizon of candidate intermediate questions, selects one to answer, and replans from the updated history. ReHoPER is task-agnostic, using the same generic instructions across datasets and models without labeled data or task-specific prompt design. Across multiple datasets, including iLLC, a new controlled benchmark for compositional reasoning, ReHoPER outperforms strong baselines, with the largest gains in the most compositional settings. Our implementation and the iLLC generator are publicly available to support future work.
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
ACM classes: I.2.7
Cite as: arXiv:2610.00940 [cs.CL]
  (or arXiv:2610.00940v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.00940

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Saeed Ahmadnia [view email]
[v1] Thu, 1 Oct 2026 02:16:27 UTC (364 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org