arXiv:cs.LG· Wenwen Si, Honghao Wei·· 4 小时前AI 评分37
基于共形动作集的强化学习:RLCP 及其在序列推荐中的应用
Reinforcement Learning with Conformal Action Sets: An Application to Sequential Recommendation
AI 导读
研究者提出 Calibrated Pruning 强化学习(RLCP),用 critic 分数和在线阈值动态调整保留动作集,阈值由代理目标是否命中的二值反馈更新,并证明了自适应轨迹上代理漏检率的确定性上界。
正文
Abstract:Sequential recommenders typically use a fixed slate size even though the number of useful alternatives changes within a session. We propose Reinforcement Learning with Calibrated Pruning (RLCP), which adapts the retained action set using critic scores and an online threshold. The threshold is updated from binary feedback indicating whether the set contains an action in a proxy target. We prove a deterministic bound on the observed proxy miss rate along adaptive trajectories. To quantify the effect of pruning on reward, we derive an exact decomposition of value loss into filtering and selection losses. Under explicit proxy and critic approximation conditions, this decomposition yields a finite session reward bound that also accounts for imperfect selection and set truncation, without requiring the learning parameters to converge. Experiments on KuaiRand-Pure and MovieLens 1M compare two RLCP implementations with four RL baselines. In each of the 19 configurations, at least one RLCP variant achieves the highest catalog diversity, reaching $1.11\times$ to $5.21\times$ that of the strongest baseline, with competitive session depth and no larger retained sets.
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.08743 [cs.LG] |
| (or arXiv:2610.08743v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08743 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Wenwen Si [view email]
[v1]
Tue, 6 Oct 2026 17:40:52 UTC (861 KB)
来源:arXiv:cs.LG · arxiv.org