arXiv:cs.AI· Jing Ding, Ziqiao Ma, Jiayuan Mao, Joyce Chai, Freda Shi·· 4 小时前AI 评分35
Vision Transformers 抽象概念落地的转喻回路研究
Metonymic Circuits for Abstract Concept Grounding in Vision Transformers
AI 导读
研究揭示 Vision Transformers 通过转喻机制将抽象概念落地:抽象预测由具体可解释的锚概念(如 fire)驱动。团队在 CLIP 和 DINO 视觉编码器上应用 Transcoders,恢复出可关联具体概念语义标签的中间特征,并在定制图标数据集上追踪到结构化转喻回路——浅层以感知基元为主,类物体锚概念先于抽象目标出现;含渲染文字的图像则走独立的感知到文本路径。
正文
Abstract:We study how Vision Transformers ground abstract concepts (e.g., angry) when training data provide limited direct referential evidence. We hypothesize a metonymic grounding mechanism in which abstract predictions are driven by concrete, interpretable anchor concepts (e.g., fire) that bridge visual signals to abstract semantics. By applying Transcoders on CLIP and DINO vision encoders, we recover intermediate features that can be associated with semantic labels for more concrete concepts, and trace their contributions in circuits underlying abstract concept recognition. Experiments on a carefully curated icon dataset reveal structured metonymic circuits, in which perceptual primitives dominate early layers and object-like anchors precede abstract targets. Images containing rendered text instead recruit a distinct perceptual-to-textual route. Causal interventions further validate that metonymic intermediates are functionally involved in grounding abstract concepts.
| Comments: | EMNLP 2026 Main. Project Website: this https URL |
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.06928 [cs.AI] |
| (or arXiv:2610.06928v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.06928 arXiv-issued DOI via DataCite |
Submission history
From: Jing Ding [view email]
[v1]
Sat, 3 Oct 2026 00:53:32 UTC (3,894 KB)
来源:arXiv:cs.AI · arxiv.org