跳到正文
arXiv:cs.CL· Thomas Vaitses Fontanari, Maximo Eduardo Rulli, Federico Alvetreti, Donatella Genovese, Simone Scardapane·· 3 小时前AI 评分39

LRCC:用条件计算泛化低秩压缩

LRCC: Generalizing Low-Rank Compression with Conditional Computation

AI 导读

研究人员提出低秩条件计算(LRCC),在每个 Transformer 块训练一个轻量路由器,为每个 token 在少量嵌套低秩路径中选择,训练时冻结低秩因子、只优化路由器。

正文

View PDF HTML (experimental)

Abstract:Low-rank compression reduces the cost of pretrained language models by replacing linear transformations with low-rank factorizations. However, conventional methods use a fixed rank allocation during inference, assigning the same amount of compute regardless of the input token. We introduce Low-Rank Conditional Computation (LRCC), which adds token-dependent computation to pretrained models by training one lightweight router per Transformer block to select among a small set of nested low-rank paths. During training, the low-rank factors remain frozen, and only the routers are optimized. We evaluate LRCC on Llama and Qwen models for language modeling and zero-shot downstream tasks. Within the same average active-parameter budget, LRCC improves the predictive performance over static low-rank compression, including a 7.6 percentage-point gain in average downstream accuracy on Llama-2-7B over static methods. At matched batch-size-1 decoding latency, LRCC improves both perplexity and downstream accuracy on Llama-3.2-1B and remains competitive on Llama-2-7B, without specialized kernels. Finally, we assess the usefulness of assigning a token-wise path by analyzing the routers' path choices.
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2610.08858 [cs.CL]
  (or arXiv:2610.08858v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.08858

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Thomas Vaitses Fontanari [view email]
[v1] Mon, 5 Oct 2026 09:01:51 UTC (335 KB)

来源:arXiv:cs.CL · arxiv.org