跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Allen Tran, Jia Wan, Nathan Kallus, Aur\'elien Bibaut·· 14 小时前AI 评分36

可扩展的多任务逆强化学习

Scalable Multi-Task Inverse Reinforcement Learning

AI 导读

研究者提出一种多任务逆强化学习方法,在低秩假设下汇集同一环境中不同奖励的多个智能体数据,缓解覆盖要求并实现对新环境下多任务的可扩展评估,规划计算量随秩而非任务数增长。该方法在奖励恢复和新环境策略学习上给出有限样本保证,实验显示其对有限覆盖稳健,能以低于基线的遗憾值迁移到目标环境,且任务越多计算优势越明显。

正文

View PDF HTML (experimental)

Abstract:By learning transferable rewards, inverse reinforcement learning (IRL) enables counterfactual evaluation of agents under modified environments. Such transfer places strict requirements on coverage since target environments affect agents' state occupancy. We propose a multi-task IRL method that pools data across multiple agents with different rewards in the same environment under a low-rank assumption. In addition to alleviating coverage requirements, so each task need not visit every state as long as others do, the method offers scalable evaluation of multiple tasks under new environments as computationally intensive planning scales with rank rather than the number of tasks. We provide finite sample guarantees on reward recovery and on policy learning in new environments. Experiments show our method is robust to limited coverage, recovers rewards on and off of each task's support, transfers to target environments at lower regret than baselines, with its computational advantage over per-task methods widening as tasks grow.
Subjects: Machine Learning (cs.LG); Machine Learning (stat.ML)
Cite as: arXiv:2610.00758 [cs.LG]
  (or arXiv:2610.00758v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.00758

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Jia Wan [view email]
[v1] Wed, 30 Sep 2026 21:52:03 UTC (264 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org