arXiv:cs.LG(机器学习,全量分类)· Xinyu Li, Mononito Goswami, Hao Liu, Nikos Kanakaris, Langlin Huang, Prithwish Jana, Patrick Bl\"obaum, Purak Jain·· 5 小时前AI 评分50
Hermes:学习上下文推理以解锁测试时扩展
Hermes: Learning Contextual Reasoning Unlocks Test-Time Scaling
AI 导读
论文提出 Hermes,一组可配置的 harness,逐步把多上下文窗口间的分配与复用决策交给模型,并配套 Hermes-Learn 两阶段训练框架学习这种上下文推理能力。实验显示强模型能利用这种灵活性随推理时算力扩展,较小开源模型经 Hermes-Learn 训练后可缩小差距,收益可跨基准和模型泛化、外推到训练时未见算力,并能迁移到 Hermes 之外的测试时扩展方法。
正文
Abstract:Test-time scaling improves model performance by allocating additional compute during inference. Using this compute effectively across multiple context windows requires deciding how to allocate fresh contexts and what information to carry between them. We call a model's ability to make these decisions contextual reasoning. Existing approaches largely prescribe these decisions through their harness; we instead shift them to the model. We introduce 1) Hermes, a family of simple, configurable harnesses that progressively varies model control over context allocation and reuse, and 2) Hermes-Learn, a two-stage framework for learning these capabilities. We find that capable models can exploit this flexibility to scale with additional inference-time compute, while smaller open-source models initially struggle to do so. Training with Hermes-Learn closes this gap, inducing adaptive contextual reasoning strategies that vary with both the problem and the progress of reasoning. These gains generalize across benchmarks and models, extrapolate beyond the inference-time compute seen during training, and transfer to complementary test-time scaling methods beyond Hermes.
| Comments: | 45 pages |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL) |
| Cite as: | arXiv:2609.38332 [cs.LG] |
| (or arXiv:2609.38332v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2609.38332 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Mononito Goswami Dr. [view email]
[v1]
Tue, 29 Sep 2026 18:01:13 UTC (376 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org