跳到正文
arXiv:cs.AI· Shengxiang Gao, Chao Lei, Jey Han Lau, Linhao Luo, Jianzhong Qi·· 6 小时前AI 评分34

Reward on Path:为知识图谱问答学习中间监督信号

Reward on Path: Learning Intermediate Supervision Signals for Knowledge Graph Question Answering

AI 导读

针对知识图谱问答中答案标签监督将所有通向答案的路径视为正确、信号噪声大的问题,研究者提出 RoP 框架,用非对称目标从答案标签学习轻量的、以问题为条件的路径奖励。该奖励经蒸馏与 GRPO 在策略优化两阶段训练 LLM 关系路径生成器,在多个 KGQA 数据集上 F1 较答案标签方法至少提升 2.6%,且无需 LLM 精炼监督的高昂成本。

正文

View PDF HTML (experimental)

Abstract:Knowledge Graph Question Answering (KGQA) aims to answer user questions by reasoning over Knowledge Graphs (KGs). Recent methods use supervision derived from answer labels or refined by Large Language Models (LLMs) to train models that retrieve KG evidence for LLM-based answer reasoning. However, answer-derived supervision treats every answer-reaching path as correct and thus yields noisy training signals, whereas LLM-refined supervision mitigates this noise at substantial cost. To address these limitations, we propose Reward on Path (RoP), a framework to learn a lightweight, question-conditioned path reward from answer labels with an asymmetric objective. Paths reaching the same answer are supervised jointly as a bag, allowing the reward model to learn their relative contributions, while each path is penalized individually for retrieving non-answer entities. The learned reward then trains an LLM-based relation path generator in two stages: reward distillation transfers reward-induced preferences over candidate paths into the generator, and on-policy optimization with GRPO further refines the policy on self-generated paths. Generated paths are grounded in the KG to retrieve evidence for answer reasoning. Experiments on multiple KGQA datasets show that RoP improves F1 over answer-derived methods by at least 2.6\% while outperforming LLM-refined methods without costly supervision construction.
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2605.10791 [cs.AI]
  (or arXiv:2605.10791v2 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2605.10791

arXiv-issued DOI via DataCite

Submission history

From: Shengxiang Gao [view email]
[v1] Mon, 11 May 2026 16:21:47 UTC (275 KB)
[v2] Thu, 1 Oct 2026 15:40:30 UTC (496 KB)

来源:arXiv:cs.AI · arxiv.org