跳到正文
arXiv:cs.CL· Hanzuo Liu, Chunyu Liu, Chaofan Lin, Alex Lamb, Mingyu Gao·· 3 小时前AI 评分37

EncBank:在 LLM 查询间复用编码器,实现紧凑可复用记忆

Cache the Encoder Within:Compact, Reusable Memory across LLM Queries

AI 导读

EncBank 将预训练 LLM 的下层作为可复用文档编码器,紧凑存储其输出供适配后的上层读取器使用。在三个 Qwen 主干、五个基准套件上,4-bit 存储使各基准总分与原生精度差距在 1 分以内,并在 Qwen3-8B 负载中仅保留 28.1% 的持久 GPU 存储占用。

正文

View PDF HTML (experimental)

Abstract:Repeated queries over shared documents incur redundant encoding, while caching model states introduces persistent storage costs. Building on CoMem's intermediate-state interface, EncBank treats a pretrained LLM's lower layers as a reusable document encoder and compactly stores their outputs for an adapted upper-layer reader. A self-distilled suffix adapter is shared across storage precisions within each backbone, without quantization-specific retraining. Across five benchmark suites on three Qwen backbones spanning different sizes and full-attention and hybrid architectures, 4-bit storage keeps each reported benchmark aggregate within one score point of native-precision EncBank. In a fixed Qwen3-8B workload, it retains 28.1% of the native-precision persistent GPU store. Separate native-precision controls yield a 1.40x selected-pack prefill speedup over same-evidence, same-adapter text replay, at a 3.12-point RULER accuracy cost. A native-precision Qwen3.8-27B configuration also passes 70 of 89 Terminal-Bench 2.1 tasks. EncBank thus combines reusable computation with compact memory, while task fidelity and end-to-end benefits remain dependent on the workload, preparation costs, and reuse frequency.
Comments: 17 pages, 3 figures, 7 tables
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2610.10058 [cs.CL]
  (or arXiv:2610.10058v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.10058

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Liu Hanzuo [view email]
[v1] Wed, 7 Oct 2026 13:28:25 UTC (89 KB)

来源:arXiv:cs.CL · arxiv.org