arXiv:cs.CL· Hanzuo Liu, Chunyu Liu, Chaofan Lin, Alex Lamb, Mingyu Gao·· 3 小时前AI 评分37
EncBank:在 LLM 查询间复用编码器,实现紧凑可复用记忆
Cache the Encoder Within:Compact, Reusable Memory across LLM Queries
AI 导读
EncBank 将预训练 LLM 的下层作为可复用文档编码器,紧凑存储其输出供适配后的上层读取器使用。在三个 Qwen 主干、五个基准套件上,4-bit 存储使各基准总分与原生精度差距在 1 分以内,并在 Qwen3-8B 负载中仅保留 28.1% 的持久 GPU 存储占用。
正文
Abstract:Repeated queries over shared documents incur redundant encoding, while caching model states introduces persistent storage costs. Building on CoMem's intermediate-state interface, EncBank treats a pretrained LLM's lower layers as a reusable document encoder and compactly stores their outputs for an adapted upper-layer reader. A self-distilled suffix adapter is shared across storage precisions within each backbone, without quantization-specific retraining. Across five benchmark suites on three Qwen backbones spanning different sizes and full-attention and hybrid architectures, 4-bit storage keeps each reported benchmark aggregate within one score point of native-precision EncBank. In a fixed Qwen3-8B workload, it retains 28.1% of the native-precision persistent GPU store. Separate native-precision controls yield a 1.40x selected-pack prefill speedup over same-evidence, same-adapter text replay, at a 3.12-point RULER accuracy cost. A native-precision Qwen3.8-27B configuration also passes 70 of 89 Terminal-Bench 2.1 tasks. EncBank thus combines reusable computation with compact memory, while task fidelity and end-to-end benefits remain dependent on the workload, preparation costs, and reuse frequency.
| Comments: | 17 pages, 3 figures, 7 tables |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.10058 [cs.CL] |
| (or arXiv:2610.10058v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10058 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Liu Hanzuo [view email]
[v1]
Wed, 7 Oct 2026 13:28:25 UTC (89 KB)
来源:arXiv:cs.CL · arxiv.org