arXiv:cs.LG(机器学习,全量分类)· Priya Nair, Lukas Brenner, Maya Lindqvist, Daniel Whitmore, Wen-Hsuan Liu, Tom Saliencro, Amara Okonkwo, Rohan Desai·· 18 小时前AI 评分36
VANE:按更新而非 token 打分的组合式 LoRA 专家路由方法
Score the Update, Not the Token: Descent-Aligned Routing for Combinatorial LoRA Experts
AI 导读
研究者提出 VANE,一种组合式 LoRA 专家路由方法,主张路由器应给"更新"打分而非给 token 打分。VANE 用低秩罗盘预测每个 token 的下降方向,通过更新与罗盘的对齐度对全部 reader-writer 对打分,无需实际构造更新,并以加法门激活 top-k 对。
正文
Abstract:Mixture-of-LoRA-experts methods raise the capacity of low-rank adaptation by routing each token to a few low-rank experts. Nearly all of them tie one input-side factor to one output-side factor per expert, and nearly all of them route by scoring the token: the router picks experts without seeing what any of them would write. We argue that the router should score the update. To first order, adding an expert's update to a layer output lowers the loss by the inner product between that update and the negative loss gradient at the output. This usefulness is quadratic in the token, so a router that is linear in the token sees only the part of it that runs through the token mean, and routers that rank experts by the norm of their own activations never see the output factor. If each expert is split into a reader (down-projection) and a writer (up-projection), the usefulness of every reader--writer pair becomes an inner product in the shared rank-$r$ space, and all $N_AN_B$ pairs can be scored from $N_A+N_B$ vectors. We build VANE on this identity. A low-rank compass predicts the descent direction of each token. VANE scores every pair by the alignment between its update and the compass without forming any update, activates the top-$k$ pairs with additive gates, and gives every pair its exact first-order router gradient. On single-domain commonsense reasoning and a four-domain multi-task mixture with Llama-3.2-3B and Llama-3.1-8B, VANE attains the best average among twelve PEFT and MoE-LoRA baselines, by 0.9--1.1 and 1.3--1.5 points respectively, with less than half the trainable parameters of an 8-expert MoE-LoRA. Its router scores also track the measured usefulness of experts far more closely than token routers do.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.00493 [cs.LG] |
| (or arXiv:2610.00493v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00493 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Tom Saliencro [view email]
[v1]
Wed, 30 Sep 2026 18:01:01 UTC (758 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org