跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· A. Sophia Koepke, Daniil Zverev, Shiry Ginosar, Alexei A. Efros·· 15 小时前AI 评分40

研究质疑柏拉图表示假说:文本与图像表征仅共享粗粒度结构

Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale

AI 导读

一项研究重新检验柏拉图表示假说,发现支持跨模态表征收敛的证据远比后续工作所称的薄弱。原研究在 1024 对文本-图像上使用的互 k 近邻指标只能捕捉粗粒度结构,随数据规模扩大 k 须成比例增长,且对齐度随语言模型增强而饱和,一对一配对也高估了对齐程度。图像与文本表征确实共享粗粒度语义结构,但更强语言模型和更丰富描述都未带来细粒度对齐。

正文

View PDF HTML (experimental)

Abstract:The Platonic Representation Hypothesis posits that neural networks trained on different modalities (e.g., text and images) converge toward a shared representation of reality. If true, this has significant implications for whether modality choice matters at all. In this paper, we show that the evidence for this claim is substantially weaker than subsequent work suggests. The mutual $k$-nearest-neighbor metric used on 1024 text-image pairs in the original study captures only coarse structure. To keep the alignment from collapsing as one scales up the data, $k$ has to grow proportionally, undercutting the argument for fine-grained representational convergence. The reported increase in alignment with language model strength saturates for recent models. Moreover, the one-to-one text-image pairing favors alignment, while alignment decreases with non-bijective data. We further find that image and text representations indeed share coarse semantic structure, but neither stronger language models nor richer captions yield fine-grained alignment. Thus, multimodal representations share coarse structure without evidence of convergence to a shared representation -- arguably, full representational convergence would require fine-grained alignment.
Comments: Project page: this http URL
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2604.18572 [cs.CV]
  (or arXiv:2604.18572v3 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2604.18572

arXiv-issued DOI via DataCite

Submission history

From: A. Sophia Koepke [view email]
[v1] Mon, 20 Apr 2026 17:56:02 UTC (10,094 KB)
[v2] Tue, 2 Jun 2026 17:45:12 UTC (11,355 KB)
[v3] Thu, 1 Oct 2026 05:16:14 UTC (16,076 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org