arXiv:cs.AI· Peter Devine, Nick Ryan, Benjamin Sirb, Alex Chiocchi·· 3 小时前
Internalizer:面向超大规模语言模型的便携式上下文到参数映射
Internalizer: Portable Context-to-Parameter Mapping for Very Large Language Models
AI 导读
arXiv 论文提出 Internalizer,一种可移植的上下文到参数映射超网络,能为冻结的 284B 参数 DeepSeek v4 Flash 生成文档专属 LoRA 适配器,规模比此前工作大两个数量级。
正文
Abstract:Hypernetworks that map a context directly to a LoRA adapter let a large language model carry that context in its weights, but prior work has demonstrated them only on base models of up to 14 billion parameters.
We present the Internalizer, a state-of-the-art, portable Context-to-Parameter Mapping hypernetwork that generates document-specific LoRA adapters for the frozen 284B-parameter DeepSeek v4 Flash, a target two orders of magnitude larger than in any previous work. Most of its parameters live in a model-agnostic trunk with only thin entry and exit layers per base model, so it trains cheaply against small models before being ported to the large one.
On unseen documents of up to 4096 tokens, the generated adapters reach 84.9% top-1 and 97.8% top-5 teacher-forced accuracy against 63.4% and 83.5% for the base model, with nothing in the context window but a three-word instruction.
Once the hypernetwork is trained, a single forward pass turns any document into an adapter for such a model, which could be served alone for speed or alongside the document in the window to raise accuracy further.
| Comments: | 14 pages, 3 figures, 1 table |
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.11715 [cs.AI] |
| (or arXiv:2610.11715v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11715 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Peter Devine [view email]
[v1]
Thu, 8 Oct 2026 11:19:13 UTC (111 KB)
来源:arXiv:cs.AI · arxiv.org