arXiv:cs.LG· Taiheng Pan·· 3 小时前AI 评分37
Probe the Harness:语言模型陈旧数据 RL 对比中的实验装置检查
Probe the Harness: Setup Checks for Stale-Data RL Comparisons in Language Models
AI 导读
研究者提出 PTH(Probe The Harness)检查清单,指出实验装置细节可反转陈旧样本训练方法的排名。在 SAN 与截断重要性采样(TIS)于 verl 及单 GPU 训练器上的对比中,PPO 比率计算、数据种子、replay 队列复用首批数据 33 次、两个损失归一化器描述不符这四项细节改变了结论;检查后 TIS 在 verl 上与 SAN 持平。
正文
Abstract:Methods for training language models on stale samples are judged by comparisons against importance-corrected baselines. We show that details of the experimental harness can reverse the observed ranking of methods, and we introduce PTH (Probe The Harness), a set of checks that makes the harness visible. Our case is a comparison between SAN, a behaviour-free method, and truncated importance sampling (TIS) on verl and in a single-GPU trainer, in which SAN first finished ahead in both stacks. Four details of the harness changed this comparison: the PPO ratio was taken against the learner's own recomputed probabilities, the data seed did not reach the TIS arm, the replay queue reused its first batch for 33 updates, and two loss normalisers differed from their description. In each case the logged quantity looked consistent with a working setup, while the quantity that defines the comparison went unchecked. With the harness checked, TIS matches SAN on verl, and in the trainer TIS learns steadily while SAN keeps a margin. We contribute the signature of each detail and its effect on the comparison, reference results for TIS and uncorrected GRPO under sampler lag, and the PTH checklist.
| Comments: | 8 pages, 2 figures, 4 tables |
| Subjects: | Machine Learning (cs.LG); Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.02911 [cs.LG] |
| (or arXiv:2610.02911v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02911 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Taiheng Pan [view email]
[v1]
Fri, 2 Oct 2026 07:02:13 UTC (92 KB)
来源:arXiv:cs.LG · arxiv.org