arXiv:cs.LG· Miaobo Hu, Shuhao Hu, Xiaobo Guo, Xin Wang, Bokun Wang, Daren Zha, Jun Xiao·· 3 小时前AI 评分36
FSPO:面向预算受限 LLM RL 后训练的策略一致风险与 Pareto 可行控制
FSPO: Policy-Consistent Risk and Pareto-Feasible Control for Budgeted LLM RL Post-Training
AI 导读
FSPO 是一个面向预算受限 LLM RL 后训练的回馈状态控制器,联合解决风险模型与部署控制器不一致、选择性动作后校准失效、多资源可行性无法认证三个问题。在匹配 GRPO 资源包络下,FSPO 取得 66.11% 留出集与 59.43% OOD 准确率,优于最强自适应基线 PB2 的 64.47% 与 57.03%。
正文
Abstract:Adaptive LLM reinforcement-learning post-training changes multiple training actuators online, including rollout temperature, group size, clipping, KL regularization, verifier allocation, and update budget. Three coupled issues remain unresolved. A future-risk model trained from behavior trajectories need not estimate the risk induced by the controller that will be deployed; a score calibrated on logged state-action pairs can become miscalibrated after selective action choice; and independent per-resource minimum costs do not in general certify a feasible multi-resource continuation. We introduce FSPO, a feedback-state controller for budgeted LLM RL post-training that addresses these issues jointly. FSPO learns a policy-consistent risk-to-go model whose Bellman target follows the same frozen controller used for future decisions, together with a long-horizon utility model. Decision-conditioned trajectory calibration (DCTC) calibrates risk on cross-fitted trajectories generated by actions selected by provisional controllers. A Pareto resource continuation certificate (PRCC) admits an action only when a non-dominated cumulative reservation remains feasible over the residual horizon. Under a matched GRPO resource envelope, FSPO reaches 66.11% held-out and 59.43% OOD accuracy, compared with 64.47% and 57.03% for PB2, the strongest evaluated adaptive baseline. Three paired training seeds give gains of +2.42 and +3.19 percentage points over the contextual bandit on held-out and OOD evaluation. Under high behavior-deployment mismatch, policy-consistent risk lowers selected-decision ECE from 0.108 to 0.053; DCTC lowers it from 0.039 to 0.022 at matched acceptance; PRCC removes false-feasible admissions on an 18-action catalog ($0.197\rightarrow0.000$); and enabling all three components reduces trajectory failure from 0.181 to 0.083 in a factorial ablation.
| Comments: | 40 pages, 4 figures |
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.02828 [cs.AI] |
| (or arXiv:2610.02828v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02828 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Miaobo Hu [view email]
[v1]
Fri, 2 Oct 2026 05:19:54 UTC (1,426 KB)
来源:arXiv:cs.LG · arxiv.org