arXiv:cs.LG(机器学习,全量分类)· Zhanming Zhang, Vinoth Selvendran·· 14 小时前AI 评分41
无数据训练先锐化:面向测试时强化学习的进入状态锐化方法
Sharpen Before You Adapt: Data-Free Entry-State Sharpening for Test-Time Reinforcement Learning
AI 导读
研究者提出"进入状态锐化":在 TTRL 之前用无数据训练把通用 checkpoint 调整到更利于无标签自监督适配的状态。
正文
Abstract:Test-time reinforcement learning (TTRL) adapts language models on unlabeled test problems using supervision derived from their own samples. This makes the checkpoint's \emph{entry state} consequential: a diffuse policy provides noisier self-supervision and may spend much of a limited adaptation budget merely concentrating probability mass before reliably expressing capability it already possesses. We propose \textbf{entry-state sharpening}: use data-free training \emph{before} TTRL to prepare a general-purpose checkpoint in a state that subsequent label-free adaptation can exploit more efficiently. The idea is not tied to one training recipe; different data-free objectives can move the same base model to different entry states. Across five data-free checkpoints derived from Qwen3-4B and evaluated under an identical 15-step TTRL protocol, entry policy entropy strongly rank-orders endpoint conversion efficiency, a reliability-to-reachability measure (Spearman $\rho=-0.90$; $\rho=-0.99$ after controlling for entry reachability). The contrast across objectives is striking: R-Zero remains diffuse at $3.39$ nats and finishes below the untuned base in 6/6 matched comparisons across MATH, GPQA, and AMC, whereas SPIRAL reaches $0.07$ nats and achieves the highest post-TTRL accuracy on MATH and GPQA despite its self-play stage using no math training data. An in-domain label-free self-distillation intervention further shows that the entry state can be deliberately sharpened. These results motivate treating checkpoint preparation as a \emph{state-control problem}: use data-free training to improve TTRL readiness, with entry entropy as a label-free control signal and reachable capability as the constraint.
| Comments: | 10 pages, 2 figures, 3 tables |
| Subjects: | Machine Learning (cs.LG) |
| ACM classes: | I.2.6; I.2.7 |
| Cite as: | arXiv:2610.00903 [cs.LG] |
| (or arXiv:2610.00903v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00903 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Zhanming Zhang [view email]
[v1]
Thu, 1 Oct 2026 01:33:30 UTC (268 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org