arXiv:cs.CL· Jiahao Lu, Mohan Kankanhalli·· 3 小时前
Agentic-TTT:用测试时策略决定 LLM 何时做测试时训练
Agentic-TTT: Training test-time policy for test-time training
AI 导读
针对测试时训练(TTT)并非普遍有效的问题,研究者提出 Agentic-TTT,学习一个测试时策略来决定何时启用 TTT、调用哪种算法、能否复用已有技能,并将 TTT 流程封装为可调用工具。在其基准上,Agentic-TTT 相对骨干模型将效用提升近一倍,还能权衡效用与算力,并泛化到训练中未见过的领域。
正文
Abstract:Test-time training (TTT) adapts an LLM's parameters using signals derived from test inputs, and can make striking improvements in pre-specified settings such as IMO competitions or designated open problems. By turning deployment experience into parameter updates, TTT provides a direct mechanism for model-level self-improvement. Yet TTT is not universally beneficial: each TTT algorithm works in different settings, and applying an ill-suited method could waste test-time compute or even damage model performance. Therefore, such parameter-level self-improvement requires agency: the model must decide when TTT is warranted, which algorithm to invoke, and whether an existing skill can be reused. To fill this gap, we introduce Agentic-TTT, which learns a test-time policy to govern those decisions. Agentic-TTT turns TTT procedures into callable tools, treats accumulated skills as an evolving deployment environment, and trains its policy using the observed utility gains from its decisions. On our benchmark, Agentic-TTT nearly doubles the utility over the backbone model, learns to trade off utility against compute, and generalizes to domains unseen during training. Together, these results point toward autonomous self-improvement: models that can decide how to learn from their own deployment experience.
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.12002 [cs.LG] |
| (or arXiv:2610.12002v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.12002 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jiahao Lu [view email]
[v1]
Thu, 8 Oct 2026 14:06:26 UTC (821 KB)
来源:arXiv:cs.CL · arxiv.org