arXiv:cs.AI· Naveen Vakada, Mingyuan Li, Shaoxiong Ji·· 3 小时前
无标签引导:将测试时强化学习压缩至仅 bias 子空间
Label-free steering: Compressing test-time reinforcement learning into bias-only subspaces
AI 导读
研究者提出 label-free bias-only TTRL,以多数投票伪标签作为奖励,仅优化约 100K 个 bias 参数并冻结预训练骨干。在 MATH-500 上,Qwen2.5-7B 达到 76.67% 准确率,略超其有标签 bias-steering 复现结果,且优化参数比全参数 TTRL 少 76,000 倍。
正文
Abstract:Test-time reinforcement learning (TTRL) enables models to improve their reasoning without relying on labeled training data, but existing approaches typically optimize a large fraction of the model parameters. This raises a natural question: can effective test-time adaptation emerge when both the reward signal and the optimization space are severely restricted? We answer this question with label-free bias-only TTRL, which uses majority-vote pseudolabels as rewards and optimizes only ~100K bias parameters while keeping the pretrained backbone frozen. On MATH-500, our approach reaches 76.67% accuracy with Qwen2.5-7B, slightly exceeding our own labeled bias-steering reproduction while optimizing 76,000x fewer parameters than full-parameter TTRL. The same training procedure improves performance across vision-language and audio reasoning tasks, including MathVista, AI2D, LogicVista, and MMAU. We further show that the learned steering vectors transfer to 4,500 held-out MATH problems, indicating that the adaptation is not limited to the problems used during test-time optimization. Finally, we analyze why this highly restricted adaptation can work, showing that majority-vote reliability improves with rollout consensus and that bias subspaces with greater accessible gradient energy exhibit stronger downstream trainability. These results demonstrate that substantial test-time adaptation can emerge from optimizing a tiny bias-only subspace using entirely label-free rewards.
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2609.18587 [cs.LG] |
| (or arXiv:2609.18587v3 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2609.18587 arXiv-issued DOI via DataCite |
Submission history
From: Naveen Vakada [view email]
[v1]
Wed, 16 Sep 2026 12:45:25 UTC (224 KB)
[v2]
Mon, 28 Sep 2026 18:33:52 UTC (397 KB)
[v3]
Thu, 8 Oct 2026 13:17:54 UTC (608 KB)
来源:arXiv:cs.AI · arxiv.org