arXiv:cs.AI· Keerthi Kaashyap, Dennis Anthony, Akshay Krishnan, Nhi Ngoc Nguyen, Jeremy Collins, James Hays, Shreyas Kousik, Animesh Garg·· 4 小时前AI 评分41
SNAP:用更弱解码器实现更强几何表征的新视角合成方法
Less Decoder is More Encoder: Geometric Representation Learning from Novel View Synthesis
AI 导读
研究者提出 SNAP,一个自监督 encoder-decoder Transformer,通过姿态条件局部解码器和潜空间重建目标,解决现有 NVS 方法因解码器空间表达力过强、像素级目标而表征能力差的问题。
正文
Abstract:This paper examines the role of Novel View Synthesis (NVS) in geometric representation learning. In principle, NVS should reason about 3D scene structure, thereby enabling transferable multi-view geometric representations. Yet, existing encoder-based NVS methods yield poor representations. This is not because of a lack of supervisory signal, but rather due to inconspicuous architectural choices: \textit{spatially expressive decoders} that dilute representational capabilities of the scene encoder, and \textit{low-level pixel-space targets} that hinder feature learning. We present SNAP, a self-supervised encoder-decoder transformer that addresses both through a pose-conditioned local decoder and a latent-space reconstruction objective. SNAP is task agnostic, and we show that it is competitive with special-purpose geometry-supervised methods. SNAP also performs competitively against self-supervised representations across five tasks: visual localization, pose estimation, point correspondence, depth estimation, and robot manipulation. Remarkably, SNAP's patch features exhibit emergent viewpoint invariance that approaches heavily supervised models despite lower compute and data budgets. Under camera shifts where standard 2D representations collapse, SNAP degrades more gracefully, revealing that restricting decoder expressivity actively prevents the suppression of transferable geometric structure. this https URL
| Comments: | Accepted to NeurIPS 2026 |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Robotics (cs.RO) |
| Cite as: | arXiv:2610.03717 [cs.CV] |
| (or arXiv:2610.03717v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.03717 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Keerthi Kaashyap [view email]
[v1]
Fri, 2 Oct 2026 17:59:14 UTC (6,849 KB)
来源:arXiv:cs.AI · arxiv.org