arXiv:cs.LG· Adam Ousherovitch, Ambuj Tewari·· 2 天前AI 评分40
Compute Aligned Training:让训练目标对齐测试时推理策略
Compute Aligned Training: Optimizing for Test Time Inference
AI 导读
研究者提出 Compute Aligned Training,将推理策略视为基础策略上的算子,据此推导出新的损失函数,使 SFT 和 RL 的训练目标与测试时策略对齐。该方法针对常见测试时策略分别实例化了对应损失函数,实证表明其相比标准训练能显著提升测试时扩展效果。
正文
Abstract:Scaling test-time compute has emerged as a powerful mechanism for enhancing Large Language Model (LLM) performance. However, standard post-training paradigms, Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL), optimize the likelihood of individual samples under a base policy, creating a misalignment with test time procedures that rely on aggregated or filtered outputs. In this work, we propose Compute Aligned Training, which aligns training objectives with test-time strategies. By conceptualizing inference strategies as operators on the base policy, we derive new loss functions that maximize performance when said strategies are applied. We instantiate such loss functions for SFT and RL across common test time strategies. Finally, we provide empirical evidence that this training method substantially improves test time scaling over standard training.
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2604.24957 [cs.LG] |
| (or arXiv:2604.24957v3 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2604.24957 arXiv-issued DOI via DataCite |
Submission history
From: Adam Ousherovitch [view email]
[v1]
Mon, 27 Apr 2026 19:52:38 UTC (4,298 KB)
[v2]
Tue, 19 May 2026 21:17:01 UTC (4,298 KB)
[v3]
Wed, 30 Sep 2026 16:13:52 UTC (4,264 KB)
来源:arXiv:cs.LG · arxiv.org