跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Jonathan Williams, Esin Tureci Karthik R. Narasimhan·· 14 小时前AI 评分49

Adapter Thickets:拆分 RLVR 预算优于集中训练

Adapter Thickets: Splitting an RLVR Budget Beats Concentrating It

AI 导读

将 RLVR 训练预算拆分成多个 LoRA adapter(即 Adapter Thicket)比集中训练单个 adapter 更能提升多数投票准确率。

正文

View PDF HTML (experimental)

Abstract:Majority voting over sampled completions is the workhorse of test-time scaling, and reinforcement learning with verifiable rewards (RLVR) is the workhorse for making each completion better. The standard pipeline composes the two: train one policy with RLVR, then sample it many times and vote. We show that this composition is lossy. A vote can only overturn mistakes that its voters do not share, and RLVR sharpens a policy so that its samples increasingly make the same mistakes. With every method drawing exactly $160$ completions per problem, training a single LoRA adapter on the full RLVR budget raises single-sample accuracy on every model we test ($1.5$B-$8$B). Yet on three of four models it leaves the majority vote below that of the untrained base model, by up to $4.8$ points. The damage builds during training: voter errors grow steadily more correlated, and the majority vote accuracy peaks early before falling by up to $7.0$ points. The cause is concentration, not RLVR itself. We split the same data and training budget across $K$ LoRA adapters, each trained on its own random disjoint shard, and call the result an adapter thicket. Thickets out-vote the fully trained adapter in all $16$ (model, $K$) settings, and for $K{\geq}4$ they stay within $0.8$ points of the base model or above it. A single adapter stopped early, at a thicket member's step count, is a strong control that matches thickets for small $K$. For $K{\geq}8$, thickets keep more of RLVR's single-sample gain and out-vote this control in six of eight settings. The cost of concentration also grows with the number of votes: from $16$ to $160$ votes, the thicket's lead over the fully trained adapter widens from $1.3$ to $3.3$ points. When the plan is to sample and vote, an RLVR budget is better spent broad than deep.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.00991 [cs.LG]
  (or arXiv:2610.00991v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.00991

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Jonathan Williams [view email]
[v1] Thu, 1 Oct 2026 03:25:40 UTC (527 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org