arXiv:cs.AI· Prabin Kumar Rath, Omkar Patil, Nakul Gopalan·· 3 小时前
Keyframe Mnemonics:面向长时程行为克隆的自监督关键帧发现方法
Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning
AI 导读
研究者提出 Keyframe Mnemonics,一种自监督方法,通过从随机采样的历史观测中学习目标并作为关键帧选择奖励,发现信息关键观测(mnemonics),再训练以这些关键帧为条件的 BC 策略。
正文
Abstract:Behavior cloning (BC) in non-Markovian environments is a challenging problem because policies have to reason over contextual information over long horizons. Existing policy architectures rely on recurrent or attention-based mechanisms to capture long-term dependencies. However, recurrent models suffer from hidden-state collapse and gradient instability under backpropagation through time, while attention-based models are fundamentally limited by context length. To address these issues, we propose Keyframe Mnemonics, a novel self-supervised method that $\textit{discovers}$ a set of information-critical observations ($\textit{mnemonics}$) by learning an objective from randomly sampled past observations and using it as a reward for keyframe selection. We then train a BC policy that conditions on the discovered keyframes to model the action distribution. Under certain task-structure assumptions, our formulation provides context retention guarantees over an infinite horizon, while maintaining a small set of decision-relevant keyframes in the policy's working memory. We evaluate our method on synthetic memory domains, where mnemonic-conditioned BC policies achieve $100$% success rates (SR) and generalize to horizons orders of magnitude beyond training without performance degradation. Additionally, we evaluate on memory-intensive robot manipulation benchmark, achieving a $13.9$% average absolute SR improvement over the strongest baseline across $23$ tasks and retaining $80$% SR at $20\times$ longer horizons on a real robot. Code and videos are available at this https URL.
| Comments: | Accepted at NeurIPS 2026 |
| Subjects: | Artificial Intelligence (cs.AI); Robotics (cs.RO) |
| Cite as: | arXiv:2610.10857 [cs.AI] |
| (or arXiv:2610.10857v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10857 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Prabin Kumar Rath [view email]
[v1]
Wed, 7 Oct 2026 20:01:27 UTC (25,647 KB)
来源:arXiv:cs.AI · arxiv.org