arXiv:cs.LG· Baback Elmieh, Lynn Tsai, Zeman Li, Srinivas Kaza, Tiancheng Sun, Gabor Csapo, Ali Behrouz, Yuan Deng, Stephen Lombardi, Steven M. Seitz, Xuan Luo·· 4 小时前AI 评分42
在线神经时空记忆实现动态新视角合成,兼顾分钟级记忆保持与实时速度
Online Neural Space Time Memory for Dynamic Novel View Synthesis
AI 导读
研究提出 Online Neural Space Time Memory,通过解耦记忆更新与应用的频率,在动态人体场景中实现在线新视角合成,达到分钟级记忆保持与摊销实时速度的 SOTA 表现。该方法采用周期性记忆更新与逐帧记忆应用,并用跨视角注意力处理形变,同时引入辅助 Memory Loss 和 Memory Caching 机制抑制历史上下文漂移。
正文
Authors:Baback Elmieh, Lynn Tsai, Zeman Li, Srinivas Kaza, Tiancheng Sun, Gabor Csapo, Ali Behrouz, Yuan Deng, Stephen Lombardi, Steven M. Seitz, Xuan Luo
Abstract:Online novel view synthesis from multi-view streaming videos faces a fundamental trade-off: maintaining a persistent, long-horizon memory to reconstruct temporarily occluded regions while operating under strict real-time constraints. While Test-Time Training (TTT) offers a powerful memory mechanism, standard models mandate gradient-based memory updates at every frame to adapt to the changing motion in dynamic scenes. The computational cost of heavy memory updates precludes real-time application and can lead to instability over long contexts. Given that memory updates are more demanding than memory application and video content is largely redundant, we propose to decouple the frequencies of these two processes. Our approach performs periodic memory updates while applying the memory on a per-frame basis, using cross-view attention to manage deformations between the prior memory state and the current frame. To lock in the historical context, we introduce two critical mechanisms: an auxiliary Memory Loss that forces persistent internalization of the scene, and a Memory Caching strategy that regularizes active weights against catastrophic drift. Our method demonstrates state-of-the-art minute-scale memory persistence in online dynamic human scenes at amortized real-time speed.
| Comments: | 19 pages. Preprint. Project page with demos and video results: this https URL |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR); Machine Learning (cs.LG) |
| ACM classes: | I.4.5; I.3.7; I.2.10 |
| Cite as: | arXiv:2607.15271 [cs.CV] |
| (or arXiv:2607.15271v2 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2607.15271 arXiv-issued DOI via DataCite |
Submission history
From: Baback Elmieh [view email]
[v1]
Thu, 16 Jul 2026 17:58:18 UTC (7,586 KB)
[v2]
Mon, 5 Oct 2026 20:59:15 UTC (9,965 KB)
来源:arXiv:cs.LG · arxiv.org