跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Zhanming Zhang, Vinoth Selvendran·· 14 小时前AI 评分41

无数据训练先锐化:面向测试时强化学习的进入状态锐化方法

Sharpen Before You Adapt: Data-Free Entry-State Sharpening for Test-Time Reinforcement Learning

AI 导读

研究者提出"进入状态锐化":在 TTRL 之前用无数据训练把通用 checkpoint 调整到更利于无标签自监督适配的状态。

正文

View PDF HTML (experimental)

Abstract:Test-time reinforcement learning (TTRL) adapts language models on unlabeled test problems using supervision derived from their own samples. This makes the checkpoint's \emph{entry state} consequential: a diffuse policy provides noisier self-supervision and may spend much of a limited adaptation budget merely concentrating probability mass before reliably expressing capability it already possesses. We propose \textbf{entry-state sharpening}: use data-free training \emph{before} TTRL to prepare a general-purpose checkpoint in a state that subsequent label-free adaptation can exploit more efficiently. The idea is not tied to one training recipe; different data-free objectives can move the same base model to different entry states. Across five data-free checkpoints derived from Qwen3-4B and evaluated under an identical 15-step TTRL protocol, entry policy entropy strongly rank-orders endpoint conversion efficiency, a reliability-to-reachability measure (Spearman $\rho=-0.90$; $\rho=-0.99$ after controlling for entry reachability). The contrast across objectives is striking: R-Zero remains diffuse at $3.39$ nats and finishes below the untuned base in 6/6 matched comparisons across MATH, GPQA, and AMC, whereas SPIRAL reaches $0.07$ nats and achieves the highest post-TTRL accuracy on MATH and GPQA despite its self-play stage using no math training data. An in-domain label-free self-distillation intervention further shows that the entry state can be deliberately sharpened. These results motivate treating checkpoint preparation as a \emph{state-control problem}: use data-free training to improve TTRL readiness, with entry entropy as a label-free control signal and reachable capability as the constraint.
Comments: 10 pages, 2 figures, 3 tables
Subjects: Machine Learning (cs.LG)
ACM classes: I.2.6; I.2.7
Cite as: arXiv:2610.00903 [cs.LG]
  (or arXiv:2610.00903v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.00903

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Zhanming Zhang [view email]
[v1] Thu, 1 Oct 2026 01:33:30 UTC (268 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org