arXiv:cs.LG· Ran Li, Lei Chen·· 4 小时前AI 评分45
通过低秩激活引导实现参数高效决策算子
Learning to Decide, Not to Reason: Parameter-Efficient Decision Operators via Low-Rank Activation Steering
AI 导读
研究者提出一种名为决策算子的System-1模块,通过行为克隆训练,用33万参数即可匹配需130万参数和强化学习流程的算子,并将3685 token的推理压缩为6 token决策且精度不降。其秩-4变体仅需2.3万参数,为已发表最强技能算子的1/58,即可满足SearchQA需求。该方案在5项任务、3种基座模型上可迁移,技能可作为近似线性算子叠加、插值并在推理时热插拔。
正文
Abstract:Injecting skills into a frozen language model currently costs a million parameters and a reinforcement-learning pipeline. We introduce \method{}, a System-1 decision operator trained by behavior cloning that lowers this cost by roughly two orders of magnitude. The default operator uses 330K parameters to match a 1.33M-parameter operator trained with reinforcement learning, exceeds or achieve comparable performance, while collapsing 3,685-token deliberation into a 6-token decision with no loss in accuracy. A rank-4 variant with 23K parameters, 1/58 of the strongest published skill operator, suffices for SearchQA and near-suffices for LiveMath, where higher rank still helps; the same recipe transfers across five tasks and three backbones, with out-of-distribution gains persisting on LiveMath problems released months after training. The gap to prior work is trainability, and it is set jointly by initialization and architecture: the initialization of prior operators zeroes the gradient of both large factor matrices at the first optimization step, whereas our zero-initialized output projection inside a shared low-rank backbone receives a gradient immediately, which a gradient-flow probe confirms directly. The gain isn't chain-of-thought compression: 23 of 57 LiveMath points beat the base model's best-of-8 sampling, and a logit-lens probe shows the operator amplifies the answer along the model's existing late-layer pathway, not writing it earlier. Gains track the base model's headroom across 13 base--task pairs, and skills compose as approximately linear operators that can be added, interpolated, and hot-swapped at inference time. Code on this https URL.
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.06950 [cs.LG] |
| (or arXiv:2610.06950v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.06950 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Ran Li [view email]
[v1]
Sat, 3 Oct 2026 12:57:57 UTC (901 KB)
来源:arXiv:cs.LG · arxiv.org