arXiv:cs.LG· Xi Ding, Naichen Shi, Jiawei Zhang·· 4 小时前AI 评分41
STAVE:用两个向量替代上下文演示,实现结构化任务适配
Two Vectors Replace In-Context Demos: Structured Task Adaptation via Embeddings
AI 导读
研究者提出 STAVE,用两个任务特定向量替换上下文演示(demos):readout 向量更新生成答案的 token,context 向量更新其他结构 token 组,两者用有/无演示提示词上的答案标签训练。
正文
Abstract:In-context learning (ICL) adapts frozen large multimodal models (LMMs) to new tasks from a few demonstrations (demos), but re-encodes them at every query, where each demo image adds up to hundreds of visual tokens. Demo-free methods remove this cost with a compact task state. However, they add it at locations searched per task or at every decoder layer, where task parameters grow with depth. Moreover, inserted tokens or keys cannot change how the original prompt divides its attention within a layer. To address these issues, we propose Structured Task Adaptation via Embeddings (STAVE), which replaces demos with two task-specific vectors added to existing input embeddings. Specifically, a readout vector updates the answer-producing tokens and a context vector updates the other structural token groups. Both are trained with answer labels on prompts with and without demos. We justify these design choices theoretically using a first-order analysis of the loss and a margin bound. Extensive experiments on six LMMs and five large language models show that STAVE matches or outperforms state-of-the-art methods on multimodal tasks with far fewer task parameters and surpasses 15-shot ICL and prior task vectors on 18 text tasks, all at zero-shot inference cost.
| Comments: | Technical report |
| Subjects: | Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.07572 [cs.CL] |
| (or arXiv:2610.07572v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07572 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Xi Ding [view email]
[v1]
Tue, 6 Oct 2026 01:05:57 UTC (5,382 KB)
来源:arXiv:cs.LG · arxiv.org