跳到正文
arXiv:cs.LG· Zitong Huang, Gustavo Lucas Carvalho, Deqing Fu, Robin Jia·· 4 小时前AI 评分34

TOPL:面向分布偏移下忠实生成的 token 级离策略学习

Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift

AI 导读

研究者提出 Token-Level Off-Policy Labeling(TOPL),将后训练重构为 token 级正确性预测任务,让模型区分回复中的好 token 与坏 token。

正文

View PDF HTML (experimental)

Abstract:We propose Token-Level Off-Policy Labeling (TOPL), an off-policy training paradigm that reframes post-training as a token-level correctness prediction task. Our key intuition is that by training the model to distinguish good and bad tokens in a response, we naturally guide the model towards generating good tokens, while avoiding the pitfalls that come with directly training the model to generate off-policy tokens. Experiments on document summarization tasks show that TOPL achieves strong out-of-distribution generalization across 11 datasets against a diverse set of sequence-level and token-level baselines. We further demonstrate that TOPL transfers effectively to machine translation, suggesting that its benefits generalize across different faithful generation tasks. Through ablation studies, we confirm that our token-level learning signal is critical to good performance; sequence-level analogues do not confer similar benefits. Finally, we show that TOPL induces interpretable model updates: the LoRA adapters learned through TOPL function as linear classification heads and steering vectors.
Subjects: Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as: arXiv:2607.17524 [cs.CL]
  (or arXiv:2607.17524v2 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2607.17524

arXiv-issued DOI via DataCite

Submission history

From: Zitong Huang [view email]
[v1] Mon, 20 Jul 2026 03:57:22 UTC (3,278 KB)
[v2] Tue, 6 Oct 2026 02:14:35 UTC (4,751 KB)

来源:arXiv:cs.LG · arxiv.org