arXiv:cs.CL· Hadeel Al-Negheimish, Jasna Ilieva, Yoon Kim·· 3 小时前
用 Prolog 探测语言模型的长程演绎推理能力
Probing for Long-Horizon Deductive Reasoning Capabilities in Language Models with Prolog
AI 导读
研究团队构建 ProloNg 合成测试平台,用 Prolog 探测 LLM 的长程演绎推理能力,最难案例推理深度达 22、上下文长度 62k。测试覆盖 5 个前沿 LLM 家族的 8 个推理模型,发现随推理深度增加性能大幅下降,多数模型在深度超过 10 后接近随机水平。
正文
Abstract:Current frontier LLMs can theoretically process long contexts with 1M tokens or more. But to what extent can they go beyond simple retrieval and perform deeper reasoning over such long contexts? We empirically investigate long-horizon reasoning capabilities of LLMs, focusing on deductive logic expressed in Prolog. We construct ProloNg, a synthetic testbed to probe Prolog Long Reasoning, which systematically varies the complexity (reasoning depth) of problems, where the hardest case has a reasoning depth of 22 and 62k context length. We study 8 reasoning models across 5 families of frontier LLMs, and find that performance degrades substantially as reasoning depth grows, with the majority of models approaching chance beyond depth 10.
| Comments: | Findings of EMNLP 2026 |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.11592 [cs.CL] |
| (or arXiv:2610.11592v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11592 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Hadeel Al-Negheimish [view email]
[v1]
Thu, 8 Oct 2026 09:42:07 UTC (1,152 KB)
来源:arXiv:cs.CL · arxiv.org