arXiv:cs.CL· Jongwook Kim, Sangheon Yun·· 3 小时前
4-Tensor 注意力模型:预测场景下一语义状态,用于视频生成与机器人规划
4-Tensor Attention Model for Semantic Physical Reality
AI 导读
研究者提出 4-Tensor 注意力模型,用窗口内位置 (x, t) 与语义、时间上下文两条纤维,通过单个 softmax 联合归一化注意力,预测场景的下一语义状态。
正文
Abstract:We describe a 4-tensor attention model that predicts the next semantic state of a scene, for video generation and robot planning. A window of states has positions (x, t) and two fibers, a semantic fiber and a temporal-context fiber, and one softmax normalizes attention jointly over the window. Frames and an agent's situation are written as those states; the encoder, the renderer, and the planner remain outside the update. To test the update on its own, we train on ROCStories, where each window poses the same next-sentence task at the semantic layer. On the validation split, with one seed per setting, the last-sentence cross-entropy on the three matched settings is lower for the 4-tensor model than for a free-running one-dimensional transformer by 5.3% at H=2, L=2, by 2.6% at H=4, L=2, and by 2.4% at H=4, L=3. At H=4, L=2 the parameter counts are nearly the same, 172.5M and 175.9M. On the same two GPUs that 4-tensor run finished in 2.4 hours and the baseline run in 45.2 hours; the baseline is trained by free-running decoding, one sequential forward pass per target token.
| Comments: | 36 pages, 4 figures |
| Subjects: | Machine Learning (cs.LG); Computation and Language (cs.CL) |
| MSC classes: | 68T07, 68T50 |
| ACM classes: | I.2.6; I.2.7; I.2.10 |
| Cite as: | arXiv:2610.11716 [cs.LG] |
| (or arXiv:2610.11716v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11716 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jongwook Kim [view email]
[v1]
Thu, 8 Oct 2026 11:19:23 UTC (37 KB)
来源:arXiv:cs.CL · arxiv.org