跳到正文
arXiv:cs.LG· Xinfei Wang, Shanchen Pang, Chenhao Zhang, Shudong Wang, Wenhao Ji, Haiyuan Gui, Meng Han, Xiaojian Liao·· 3 小时前

DaCe-DT:面向异构任务的数据中心化离线多任务强化学习,基于自适应提示词与轨迹校正

DaCe-DT: Data-Centric Offline Multi-Task Reinforcement Learning via Adaptive Prompts and Trajectory Correction for Heterogeneous Tasks

AI 导读

针对离线多任务强化学习(Offline MTRL)中提示词长度利用低效、随机采样语义不相关及碎片化轨迹误导监督三大瓶颈,研究者提出数据中心化框架 DaCe-DT,集成长度门控提示词掩码(LGPM)、检索增强提示词构建(RAPC)与价值自适应回报校准(VARC)。

正文

View PDF HTML (experimental)

Abstract:Offline multi-task reinforcement learning (Offline MTRL) heavily depends on the quality and distribution of pre-collected data. However, existing methods mainly focus on algorithmic optimization, with less emphasis on data-level improvements to enhance learning ability and generalization performance. This paper, from a data perspective, reveals three key bottlenecks that limit Offline MTRL performance:(i) ineffective utilization of prompts length under diverse task complexities, and (ii) semantic irrelevance of randomly sampled prompt segments, (iii) misleading supervision induced by fragmented and discontinuous trajectories. To address these challenges, we propose DaCe-DT, a robust offline MTRL framework designed to be insensitive to heterogeneous task complexities and data quality, featuring length-gated prompt masking (LGPM), retrieval-augmented prompt construction (RAPC), and value-adaptive return calibration (VARC). Together, these mechanisms enable DaCe-DT to deliver data-centric prompt adaptation and trajectory refinement, resulting in robust multi-task generalization and stable policy learning amid heterogeneous offline data and tasks. Experimental results on Meta-World show that DaCe-DT consistently outperforms state-of-the-art methods, achieving an average improvement of 11.73% on optimal datasets and an improvement of 13.34% on suboptimal datasets, demonstrating its effectiveness in learning stably from imperfect data and improving overall multi-task performance.
Comments: 25 pages,NeurIPS-2026
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.11085 [cs.LG]
  (or arXiv:2610.11085v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.11085

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Xinfei Wang F [view email]
[v1] Thu, 8 Oct 2026 01:53:39 UTC (11,901 KB)

来源:arXiv:cs.LG · arxiv.org