跳到正文
arXiv:cs.LG· Ambarish Moharil·· 3 小时前AI 评分34

深度网络的时间几何:用双曲表示刻画训练动态以实现内在可解释性

Temporal Geometry of Deep Networks: Hyperbolic Representations of Training Dynamics for Intrinsic Explainability

AI 导读

该研究提出一种基于图的元学习框架,将 MLP 训练过程中的时间参数图嵌入 Poincaré 双曲空间,从而获得训练动态的动态双曲表示。方法对快照内神经元排列保持等变性、对历史快照排列保持不变性,在回归与分类任务上展现出较强的元网络性能。该论文已被 ICLR 2026 接收为主会议论文。

正文

View PDF HTML (experimental)

Abstract:Intrinsic explainability remains a challenging problem, particularly in contexts where multilayer perceptrons (MLPs) require dynamic re-training within an optimization environment. This paper investigates how MLPs and their training dynamics can be represented and studied in non-Euclidean spaces; our representation features the Poincaré model of hyperbolic geometry. We aim to capture the geometric evolution of their weighted topology and self-organization over time. Instead of restricting the analysis to single checkpoints---as per established measure-based explainability methods---we construct temporal \textit{parameter graphs}, i.e., snapshots over time $T$ steps of the optimization/training process for MLPs. This reflects the view that neural networks encode information not only in their weights but also in the trajectory traced during training. Drawing on the idea that many complex networks admit embeddings in hidden metric spaces where distances correspond to connection likelihood, we present a geometric and temporal graph-based metalearning framework for obtaining dynamic hyperbolic representations of the underlying neural parameter graphs. Our model embeds temporal parameter graphs in the Poincaré model ball, and learns from them while maintaining equivariance to within-snapshot neuron permutations and invariance to permutations of past snapshots. In doing so, the approach preserves functional equivalence over time and recovers the latent evolving geometry of the network. Experiments on regression and classification tasks with trained MLPs show strong meta-network performance, accompanied by hyperbolic temporal representations. This reveals how the network structure emerges over time under specific training environments, thus providing insights into the network's self-organization.
Comments: Published as a main conference paper at ICLR 2026. 10 Main Pages, 22 pages of supplementary material
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Neural and Evolutionary Computing (cs.NE); Data Analysis, Statistics and Probability (physics.data-an)
Cite as: arXiv:2610.03000 [cs.LG]
  (or arXiv:2610.03000v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.03000

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Ambarish Moharil [view email]
[v1] Fri, 2 Oct 2026 08:32:13 UTC (21,308 KB)

来源:arXiv:cs.LG · arxiv.org