arXiv:cs.LG· Ruijin Hua, Zichuan Liu, Zhuokai Zhao, Yujia Zheng·· 3 小时前
SplitJEPA:无需重建即可学习不变与变化隐世界的 JEPA
SplitJEPA: Learning Invariant and Variant Latent Worlds without Reconstruction
AI 导读
SplitJEPA 是一种无需重建、直接在表示空间联合恢复隐状态及其不变与变化结构的 JEPA 架构。论文证明,在平稳高斯预测动态与满秩变化条件下,SplitJEPA 可将不变子空间与变化子空间辨识至独立的块级等距变换,且无需引入观测解码器。合成非线性系统与机器人操作任务的实验支持了上述理论结果,并展示其在鲁棒性与效率上的实用价值。
正文
Abstract:Understanding a dynamical world calls for more than a latent state that summarizes its observations: the state should also be organized into the factors that stay shared across related observations and the factors that vary between them. For example, a robot pushing a cube to a goal should take the same action when the camera shifts or the lights dim, since nothing in the scene has moved. Existing approaches to this decomposition commonly obtain it through reconstruction, so the latent variables must first explain the entire observational world before their organization can be trusted. Joint embedding predictive architectures (JEPAs) model the latent state directly and never reconstruct, yet no existing result recovers the invariant and variant parts of the state they learn. How to learn the invariant-variant structure of the latent world without paying for its reconstruction therefore remains open. To close this gap, we introduce SplitJEPA, a JEPA that jointly recovers the latent state and its invariant and variant organization directly in representation space, without any reconstruction. We prove that, under stationary Gaussian predictive dynamics and a full-rank variation condition, SplitJEPA identifies the invariant and variant subspaces up to independent block-wise isometries, without introducing an observation decoder. Since the guarantee needs no decoder, the result extends reconstruction-free latent recovery to invariant-variant block identification. Experiments on synthetic nonlinear systems and robotic manipulation tasks support the theoretical results and show their practical value for both robustness and efficiency.
| Comments: | 23 pages, 15 figures, 6 tables |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.12349 [cs.LG] |
| (or arXiv:2610.12349v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.12349 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Ruijin Hua [view email]
[v1]
Thu, 8 Oct 2026 17:19:10 UTC (1,158 KB)
来源:arXiv:cs.LG · arxiv.org