arXiv:cs.LG· Xing Lei, Wenyan Yang, Xuetao Zhang, Donglin Wang·· 5 小时前AI 评分34
DAGR:通过差异感知目标交叉注意力实现状态条件化的目标表示
DAGR: State-Conditioned Goal Representations via Difference-Aware Goal Cross-Attention
AI 导读
针对目标条件强化学习中现有编码器均不感知当前状态、无法标记目标中仍需行动部分的问题,研究者提出 DAGR,通过多尺度门控交叉注意力将任意 late-fusion 编码器的静态嵌入细化为状态条件化表示,并以差异感知注意力按逐 token 的状态-目标失配调整分数。在 OGBench 上 DAGR 提升了导航任务表现,其余任务持平或略低;消融显示增益主要来自门控残差而非方法同名的差异偏置。代码已开源。
正文
Abstract:Goal-conditioned reinforcement learning hinges on how the goal is encoded. Contrastive, metric, temporal-distance and information-theoretic encoders disagree on the objective. They agree on one thing. None of them sees the current state, so the embedding cannot mark which part of the goal still needs action, and the policy must recover that cue by inverting both encoders. We propose DAGR, which refines the static embedding of any late-fusion encoder into a state-conditioned one through multi-scale gated cross-attention. A gated residual holds the refinement near the base, and a difference-aware attention rule biases the scores by a per-token state-goal mismatch. A single condition decides what such a refinement can guarantee, namely whether the block returns its input at closed gates. We prove that the usual post-norm placement violates it, measure the consequence on frozen checkpoints, and recover part of the resulting loss by restoring the condition. On OGBench DAGR improves navigation and matches or trails the base elsewhere. Our ablations trace the gain to the gated residual rather than to the difference bias that names the method. Code is available at this https URL
| Subjects: | Machine Learning (cs.LG); Machine Learning (stat.ML) |
| Cite as: | arXiv:2607.13731 [cs.LG] |
| (or arXiv:2607.13731v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2607.13731 arXiv-issued DOI via DataCite |
Submission history
From: Xing Lei [view email]
[v1]
Wed, 15 Jul 2026 11:45:31 UTC (3,446 KB)
[v2]
Fri, 2 Oct 2026 12:21:36 UTC (3,450 KB)
来源:arXiv:cs.LG · arxiv.org