跳到正文
arXiv:cs.CL· Jongwook Kim, Sangheon Yun·· 3 小时前

4-Tensor 注意力模型:预测场景下一语义状态,用于视频生成与机器人规划

4-Tensor Attention Model for Semantic Physical Reality

AI 导读

研究者提出 4-Tensor 注意力模型,用窗口内位置 (x, t) 与语义、时间上下文两条纤维,通过单个 softmax 联合归一化注意力,预测场景的下一语义状态。

正文

View PDF HTML (experimental)

Abstract:We describe a 4-tensor attention model that predicts the next semantic state of a scene, for video generation and robot planning. A window of states has positions (x, t) and two fibers, a semantic fiber and a temporal-context fiber, and one softmax normalizes attention jointly over the window. Frames and an agent's situation are written as those states; the encoder, the renderer, and the planner remain outside the update. To test the update on its own, we train on ROCStories, where each window poses the same next-sentence task at the semantic layer. On the validation split, with one seed per setting, the last-sentence cross-entropy on the three matched settings is lower for the 4-tensor model than for a free-running one-dimensional transformer by 5.3% at H=2, L=2, by 2.6% at H=4, L=2, and by 2.4% at H=4, L=3. At H=4, L=2 the parameter counts are nearly the same, 172.5M and 175.9M. On the same two GPUs that 4-tensor run finished in 2.4 hours and the baseline run in 45.2 hours; the baseline is trained by free-running decoding, one sequential forward pass per target token.
Comments: 36 pages, 4 figures
Subjects: Machine Learning (cs.LG); Computation and Language (cs.CL)
MSC classes: 68T07, 68T50
ACM classes: I.2.6; I.2.7; I.2.10
Cite as: arXiv:2610.11716 [cs.LG]
  (or arXiv:2610.11716v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.11716

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Jongwook Kim [view email]
[v1] Thu, 8 Oct 2026 11:19:23 UTC (37 KB)

来源:arXiv:cs.CL · arxiv.org