跳到正文
原文
HuggingFace Daily Papers(社区热门论文)·· 10 小时前AI 评分38

Memorizon:让世界模型训练突破上下文窗口限制

Memorizon: Training World Models Beyond Their Context Window

AI 导读

Memorizon 提出一种训练方法,使流式世界模型能在超出上下文窗口的长跨度上接受监督:训练样本可覆盖任意长度跨度,但只对最后 k 个 chunk 计算损失,每个被评分 chunk 通过相机共视性检索各自的 top-K latent 并组成共享 bank,bank 上限为 kK,使序列长度保持有界。

正文

Published on Sep 30

Authors:

,

,

Abstract

Streaming world models should render a place consistently across repeated visits. Directly supervising such revisits requires training samples that capture both visits, often spanning minutes. Yet dense attention over the full span incurs quadratic costs, making long-span supervision expensive. Memorizon breaks this coupling: long spans are needed for supervision, but not for attention, since the two visits can share a forward pass without including every intervening frame. A training sample covers a span of any length but is scored only on its last k chunks. Instead of tokenizing the history before them, each scored chunk retrieves its own top-K latents by camera co-visibility, and the union of these requests forms a shared bank. The bank is bounded by kK, so the sequence stays bounded however long the span; at the shortest span the recipe is exactly conventional training. Adding the bank raises the cost of a step once; beyond that, a longer span costs little, and going from 100 to 400 s adds 12% to the step time. Against a sliding-window baseline, retrieval raises revisit consistency on every split, and a span long enough to reach the first visit of each return adds a further 24% to 30%, at some cost in image quality; beyond that span, more length no longer helps. Filling the bank from another episode lowers revisit correlation by 83%, so the model uses what it retrieves. Project page: https://tingtingliao.github.io/memorizon

View arXiv page View PDF Project page GitHub Add to collection

Get this paper in your agent:

hf papers read 2610.00544

Don't have the latest CLI?

curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 1

Luffuly/memorizon

Image-to-Video •

5B •

Updated 36 minutes ago

•

27

•

1

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2610.00544 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2610.00544 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.

来源:HuggingFace Daily Papers(社区热门论文) · huggingface.co