arXiv:cs.CL· Thomas Vaitses Fontanari, Maximo Eduardo Rulli, Federico Alvetreti, Donatella Genovese, Simone Scardapane·· 3 小时前AI 评分39
LRCC:用条件计算泛化低秩压缩
LRCC: Generalizing Low-Rank Compression with Conditional Computation
AI 导读
研究人员提出低秩条件计算(LRCC),在每个 Transformer 块训练一个轻量路由器,为每个 token 在少量嵌套低秩路径中选择,训练时冻结低秩因子、只优化路由器。
正文
Abstract:Low-rank compression reduces the cost of pretrained language models by replacing linear transformations with low-rank factorizations. However, conventional methods use a fixed rank allocation during inference, assigning the same amount of compute regardless of the input token. We introduce Low-Rank Conditional Computation (LRCC), which adds token-dependent computation to pretrained models by training one lightweight router per Transformer block to select among a small set of nested low-rank paths. During training, the low-rank factors remain frozen, and only the routers are optimized. We evaluate LRCC on Llama and Qwen models for language modeling and zero-shot downstream tasks. Within the same average active-parameter budget, LRCC improves the predictive performance over static low-rank compression, including a 7.6 percentage-point gain in average downstream accuracy on Llama-2-7B over static methods. At matched batch-size-1 decoding latency, LRCC improves both perplexity and downstream accuracy on Llama-3.2-1B and remains competitive on Llama-2-7B, without specialized kernels. Finally, we assess the usefulness of assigning a token-wise path by analyzing the routers' path choices.
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.08858 [cs.CL] |
| (or arXiv:2610.08858v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08858 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Thomas Vaitses Fontanari [view email]
[v1]
Mon, 5 Oct 2026 09:01:51 UTC (335 KB)
来源:arXiv:cs.CL · arxiv.org