跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Guangzhi Xiong, Zhenghao He, Bohan Liu, Sanchit Sinha, Wenqian Ye, Aidong Zhang·· 9 小时前AI 评分42

用任务算子捕捉上下文学习动态:TO 缩小零样本推理与 ICL 差距

Capturing In-Context Learning Dynamics with Task Operators

AI 导读

研究者提出 Task Operator(TO),通过将每个注意力头的输出视为其上下文掩码对应项的仿射变换,把该变换作为解析推导的更新回放到注意力输出投影上,从而压缩上下文学习(ICL)。在词汇、算法和推理任务上,TO 取得优于既有方法的整体表现,显著缩小零样本推理与 ICL 的差距。提取的知识集中于跨层和位置的特定任务稀疏回路,对不相交演示批次取平均算子还可实现多示例扩展而无需扩大上下文窗口。

正文

View PDF HTML (experimental)

Abstract:In-context learning (ICL) enables language models to perform new tasks from demonstrations without weight updates. However, every ICL inference requires processing the full set of examples, resulting in inefficient deployments, and how ICL works mechanistically is not fully understood. Prior work compresses ICL into fixed activation vectors extracted from specific layers or positions, but these input-independent interventions fail on complex tasks where the output depends on fine-grained interactions with the input. By analyzing the ICL forward pass, we show that each attention head's output is an affine transformation of its context-masked counterpart, and that the parameters of this transformation are empirically stable across samples for a given task. Building on this, we introduce Task Operator (TO), which replays this transformation as an analytically derived update to the attention output projection. Across lexical, algorithmic, and reasoning tasks, TO achieves the best overall performance among prior methods and substantially narrows the gap between zero-shot inference and ICL. We further show that the extracted knowledge concentrates in a task-specific sparse circuit across layers and positions, and that averaging operators from disjoint demonstration batches enables effective many-shot scaling without expanding the context window. Our code is available at this https URL.
Comments: NeurIPS 2026
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2610.01054 [cs.CL]
  (or arXiv:2610.01054v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.01054

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Guangzhi Xiong [view email]
[v1] Thu, 1 Oct 2026 04:50:17 UTC (145 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org