arXiv:cs.LG(机器学习,全量分类)· Changliang Zhou, Yuanyao Chen, Rongsheng Chen, Zhiyun Lin, Zhenkun Wang·· 14 小时前AI 评分32
MiLoop:面向神经组合优化的选择性记忆传播框架
MiLoop: Selective Memory Propagation for Neural Combinatorial Optimization
AI 导读
研究者提出纯 RL 构造式框架 MiLoop,利用 rollout 中已有的多步计算进行选择性记忆传播,让浅层策略无需外部解标签或训练时搜索空间剪枝即可学到有效动态嵌入。该方法在注意力层前融合当前嵌入与历史记忆,之后施加自适应门控更新,更新后的表示同时支持当前决策与逐步复用。在四类组合优化问题、100 至 1000 万节点实例上,MiLoop 均能稳定产出高质量解,展现出强泛化能力。
正文
Abstract:Constructive neural combinatorial optimization (NCO) has emerged as a promising paradigm that learns to construct solutions to combinatorial optimization problems (COPs) step by step, which reduces reliance on handcrafted rules and enables fast inference. While many methods with dynamic embeddings generalize well, they typically rebuild subproblem representations from scratch at each step using deep attention stacks. Many high-performing methods in this category rely on solution labels or pseudo-labels for efficient training, or on aggressive search space pruning during reinforcement learning (RL). To address these limitations, we propose Memory-in-the-Loop (MiLoop), a purely RL-based constructive framework that leverages the multi-step computation already required by a rollout for selective memory propagation. Each rollout provides solution-quality feedback for learning while propagating historical representations, thereby enabling a shallow policy to learn effective dynamic embeddings without external solution labels or training-time search-space pruning. Specifically, MiLoop fuses current embeddings with historical memory before the attention layers and applies adaptive gated updates afterward. The updated representations support both current decisions and stepwise reuse. Extensive experiments across four COPs demonstrate that MiLoop consistently produces high-quality solutions on instances ranging from 100 to 10 million nodes, highlighting its strong generalization ability.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.01685 [cs.LG] |
| (or arXiv:2610.01685v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.01685 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Changliang Zhou [view email]
[v1]
Thu, 1 Oct 2026 13:42:59 UTC (4,497 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org