跳到正文
arXiv:cs.AI· Jia-Dong Zhang·· 6 小时前AI 评分30

VALSE:面向大语言模型高效推理的垂直自适应层跳过方法

VALSE: Vertical Adaptive Layer Skipping for Efficient Inference in Large Language Models

AI 导读

研究者提出 VALSE(Vertical Adaptive Layer Skipping for Efficiency),一种按样本、非连续的层跳过方法:由轻量难度估计器依据前几层为每个输入打分,逐层门控可跳过冗余的中间层并保留更深层,仅为每个输入激活必要深度,其可行性已在原型规模上初步评估。

正文

View PDF HTML (experimental)

Abstract:This paper establishes a theoretical framework for vertical adaptive layer skipping, proving three foundational results: (i) an Expected FLOPs formula (theorem 2) giving a closed-form expression for the computational cost of arbitrary per-sample skip schedules as a function of layer-wise skip probabilities; (ii) function-space superset (theorem 10) and strict inclusion (theorem 11) theorems showing that skip-layer models are strictly contained in---yet meaningfully approximate---the full-layer function space, with an explicit separating example; and (iii) a structural duality between VALSE and Mixture-of-Experts architectures (proposition 6), positioning vertical depth-wise sparsity as the orthogonal counterpart to horizontal width-wise sparsity. Building on this theory, we propose VALSE (Vertical Adaptive Layer Skipping for Efficiency), a per-sample, non-contiguous layer skipping method: a lightweight difficulty estimator scores each input from the first few layers, and per-layer gates selectively skip redundant layers---including arbitrary middle layers while retaining deeper ones---so that only the necessary depth is activated for each input, whose feasibility is preliminarily assessed at prototype scale.
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.07606 [cs.AI]
  (or arXiv:2610.07606v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.07606

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Jia-Dong Zhang [view email]
[v1] Tue, 6 Oct 2026 01:47:40 UTC (46 KB)

来源:arXiv:cs.AI · arxiv.org