跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Naveen Mysore·· 14 小时前AI 评分39

注意力流形:通过编辑学习到的 B 样条曲面来引导或阻断语言模型

Attention Manifolds: Steering or Blocking Language Models by Editing Learned B-Spline Surfaces

AI 导读

研究提出"注意力流形"——基于 query-key 交互调制 value 各维度的可学习 2D B 样条曲面,以零初始化保持预训练行为。应用于 LLaMA 3.2-1B-Instruct 和 3B-Instruct 后,WikiText-2 验证困惑度降低 2–2.5 点,参数开销仅 0.3%。

正文

View PDF HTML (experimental)

Abstract:In standard transformer attention, a source token sends the same value vector to every receiver. The query determines \emph{how much} to attend but not \emph{what} to extract. This work introduces \textbf{attention manifolds}: learned 2D B-spline surfaces $S_d(q_d, k_d)$ that modulate each value dimension based on the query-key interaction. Each surface is a tensor-product cubic B-spline initialized to zero, preserving pretrained behavior. Applied to LLaMA 3.2-1B-Instruct and 3B-Instruct, attention manifolds reduce WikiText-2 validation perplexity by 2--2.5 points with 0.3\% parameter overhead. Across 112 diverse prompts, surfaces change greedy-decoded output for 69\% (1B) to 83\% (3B) of cases, with the strongest effects on ambiguous and polysemous inputs (94--100\% change rate). The surfaces improve output quality: correcting factual errors (\emph{``the CAP theorem has three main components''} $\to$ \emph{``it is impossible to guarantee all three''}), increasing precision (\emph{``impossible to know certain properties''} $\to$ \emph{``impossible to know both position and momentum''}), and adding specificity (a generic quote $\to$ an attributed Saint Augustine citation, consistently at both scales). The learned surfaces are also mechanically editable: inverting a layer's coefficients changes greedy output for 9/10 prompts (KL~0.010), providing a geometric mechanism for model steering. Setting surface coefficients to $-1$ creates ``attention walls'' that block value flow through specific dimensions. In a preliminary experiment, a layer-wide wall redirects an explosive-device prompt from specific instructions to general educational content, suggesting a path toward safety-oriented manifold shaping.
Subjects: Machine Learning (cs.LG)
MSC classes: 68T07, 41A15
ACM classes: I.2.7; G.1.1
Cite as: arXiv:2610.00257 [cs.LG]
  (or arXiv:2610.00257v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.00257

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Naveen Mysore [view email]
[v1] Thu, 24 Sep 2026 03:55:39 UTC (918 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org