跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Tido Specht, Elias Benedict Krey, Nils Neukirch, Nils Strodthoff·· 14 小时前AI 评分37

超越线性概念:在大语言模型中发掘并对齐非线性概念流形

Beyond Linear Concepts: Discovering and Aligning Non-Linear Concept Manifolds in Large Language Models

AI 导读

研究者将计算机视觉中的 NLMCD 方法迁移到 LLM token 级激活,把概念建模为低维流形,并提出基于概念的 CBA 对齐分数。分析发现中间层与后层存在两个块结构,概念组成在网络大部分区域由句法主导,后层才转向句法与语义混合。

正文

View PDF HTML (experimental)

Abstract:Understanding information processing in large language models (LLMs) requires dissecting the geometric organization of their internal token representations. While existing mechanistic interpretability (MI) methods seek to extract concepts, they are constrained by a strong linearity assumption challenged by evidence of non-linear feature manifolds. We move beyond linear concepts by adapting Non-Linear Multi-Dimensional Concept Discovery (NLMCD) from computer vision to token-level LLM activations, modeling concepts as low-dimensional manifolds. To compare concept manifolds across layers and models, we introduce a concept-based alignment (CBA) score, a generalized Rand index that measures geometric proximity without explicit feature matching. Our analysis yields six key findings: (i) a neighboring-layer sanity check shows CBA is more sensitive than PCA- or CKA-based linear baselines; (ii) layer-by-layer alignment matrices reveal two block structures in intermediate and late layers, consistent across models and obscured by linear metrics; (iii) concept composition remains syntax-dominated through most of the network before giving way to increasingly mixed syntactic-semantic concepts in later layers, with increasing output-orientation toward the final layers; (iv) multilingual concept sharing between English and Mandarin is training-dependent rather than universal, strongest in Qwen, weaker in Llama, and absent in GPT-2; (v) inter-model alignment mirrors this structure, with strong correspondence between same-family Qwen models of different scale but weak alignment across model families; and (vi) across Tulu-3 training stages, alignment is highest between adjacent stages, with the largest shift between the base model and SFT, while subsequent preference-alignment stages (DPO, RLVR) leave early layers largely unchanged and RLVR mostly preserves DPO's concepts in late layers.
Comments: 24 pages, 13 figures. Code: this https URL
Subjects: Machine Learning (cs.LG); Computation and Language (cs.CL)
Cite as: arXiv:2610.01821 [cs.LG]
  (or arXiv:2610.01821v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.01821

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Tido Specht [view email]
[v1] Thu, 1 Oct 2026 14:55:03 UTC (4,147 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org