跳到正文
arXiv:cs.CL· Tiezheng Yu, Yuxin Jiang, Jinpeng Li, Shuning Sun, Fei Mi, Haoli Bai, Lifeng Shang·· 4 小时前AI 评分39

HARPO:面向忠实且富有创造性的语言生成的幻觉感知强化学习框架

HARPO: Hallucination-Aware Reinforcement Learning for Faithful and Creative Language Generation

AI 导读

研究者提出 HARPO 强化学习框架,通过幻觉感知生成奖励模型(HA-GRM)与选择性激活机制(SAM)联合优化语言生成的忠实性与创造性。基于 Qwen3-4B 的 HA-GRM 在 RAGTruth 上取得 78.08% 的响应级 F1,高于监督微调基线的 66.37%。

正文

View PDF HTML (experimental)

Abstract:Large Language Models (LLMs) are prone to generating hallucinated content, which compromises their reliability in knowledge-intensive tasks. To address this challenge without sacrificing creativity, we propose HARPO, a reinforcement learning framework designed to jointly optimize faithfulness and creativity. HARPO incorporates a Hallucination-Aware Generative Reward Model (HA-GRM), trained via verifiable feedback, to assess both faithfulness and writing quality. A Selective Activation Mechanism (SAM) activates writing rewards only for outputs judged hallucination-free by HA-GRM, while a data curriculum progressively shifts training from creative writing to hallucination-centric tasks. On RAGTruth, our Qwen3-4B-based HA-GRM achieves a response-level F1 score of 78.08%, compared with 66.37% for the supervised fine-tuning baseline. Experiments on Qwen2.5 and Qwen3 models from 1.7B to 8B parameters show improvements in both faithful generation and writing quality. On Qwen3-4B, HARPO reduces the HA-GRM-judged hallucination rate on MultiHopRAG from 3.29% to 1.02%, while increasing the Arena-Hard-v2.0 creative-writing score from 16.95% to 27.54%.
Comments: 11 pages
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2610.03063 [cs.CL]
  (or arXiv:2610.03063v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.03063

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Tiezheng Yu [view email]
[v1] Fri, 2 Oct 2026 09:46:40 UTC (1,283 KB)

来源:arXiv:cs.CL · arxiv.org