arXiv:cs.LG· Miaobo Hu, Shuhao Hu, Xiaobo Guo, Xin Wang, Bokun Wang, Tianshu Fu, Daren Zha, Jun Xiao·· 2 天前AI 评分35
FAER:面向语言模型后训练的可审计效用对齐轨迹回放框架
FAER: Auditable Utility-Aligned Trajectory Replay for Language Model Post-Training
AI 导读
FAER 是一个可审计的全轨迹回放框架,用于语言模型后训练中的轨迹选择,其 learner-aware 选择器 FAER-UTILITY 在 GSM8K 与 Qwen2.5-1.5B-Instruct 上达到 0.6624 质量分,高于固定选择器的 0.6329、uniform 的 0.5482 和 format-feedback 的 0.6037。
正文
Abstract:Replay selectors often rank cached trajectories by format feedback, confidence, freshness, or response length, although cache-level correctness and downstream learner utility are distinct objectives. We formalize this selection-to-learning gap and introduce FAER as an auditable full-trajectory replay framework. Its training-free fixed selector is a protocol baseline; FAER-UTILITY is the learner-aware selector fitted on disjoint calibration blocks. The normalized gradient alignment is reported as a baseline, while a disposable optimizer-aware virtual update supplies a magnitude-aware utility surface. The audit contract freezes observed fields and replay traces before evaluation labels are joined. On GSM8K with Qwen2.5-1.5B-Instruct, the matched learner study reports quality 0.6329 for the fixed selector, compared with 0.5482 for uniform and 0.6037 for format-feedback under 128 updates. Metadata-only cross-fitted calibration reaches $0.6476\!\pm\!0.0139$ over eight seeds (median 0.6481; paired 95% interval $[+0.079,+0.122]$) at 63,276 target-run tokens; its recorded full cost is 189,642 tokens and 3.48 GPU-hours including calibration. The completed FAER-UTILITY row reaches 0.6624 at 62,844 target-run tokens and 4.26 GPU-hours. Format-feedback selects records with correctness 0.6953, compared with 0.3594 for the fixed selector, despite the different downstream ranking. The completed comparison surfaces report the learner-aware ablation, same-seed gap, policy-optimization rows, and strict zero-shot transfer.
| Comments: | 35 pages, 6 figures |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (stat.ML) |
| Cite as: | arXiv:2610.00385 [cs.LG] |
| (or arXiv:2610.00385v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00385 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Miaobo Hu [view email]
[v1]
Wed, 30 Sep 2026 09:41:08 UTC (1,620 KB)
来源:arXiv:cs.LG · arxiv.org