arXiv:cs.CL· Tiezheng Yu, Yuxin Jiang, Jinpeng Li, Shuning Sun, Fei Mi, Haoli Bai, Lifeng Shang·· 4 小时前AI 评分39
HARPO:面向忠实且富有创造性的语言生成的幻觉感知强化学习框架
HARPO: Hallucination-Aware Reinforcement Learning for Faithful and Creative Language Generation
AI 导读
研究者提出 HARPO 强化学习框架,通过幻觉感知生成奖励模型(HA-GRM)与选择性激活机制(SAM)联合优化语言生成的忠实性与创造性。基于 Qwen3-4B 的 HA-GRM 在 RAGTruth 上取得 78.08% 的响应级 F1,高于监督微调基线的 66.37%。
正文
Abstract:Large Language Models (LLMs) are prone to generating hallucinated content, which compromises their reliability in knowledge-intensive tasks. To address this challenge without sacrificing creativity, we propose HARPO, a reinforcement learning framework designed to jointly optimize faithfulness and creativity. HARPO incorporates a Hallucination-Aware Generative Reward Model (HA-GRM), trained via verifiable feedback, to assess both faithfulness and writing quality. A Selective Activation Mechanism (SAM) activates writing rewards only for outputs judged hallucination-free by HA-GRM, while a data curriculum progressively shifts training from creative writing to hallucination-centric tasks. On RAGTruth, our Qwen3-4B-based HA-GRM achieves a response-level F1 score of 78.08%, compared with 66.37% for the supervised fine-tuning baseline. Experiments on Qwen2.5 and Qwen3 models from 1.7B to 8B parameters show improvements in both faithful generation and writing quality. On Qwen3-4B, HARPO reduces the HA-GRM-judged hallucination rate on MultiHopRAG from 3.29% to 1.02%, while increasing the Arena-Hard-v2.0 creative-writing score from 16.95% to 27.54%.
| Comments: | 11 pages |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.03063 [cs.CL] |
| (or arXiv:2610.03063v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.03063 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Tiezheng Yu [view email]
[v1]
Fri, 2 Oct 2026 09:46:40 UTC (1,283 KB)
来源:arXiv:cs.CL · arxiv.org