arXiv:cs.LG· Zhouyu Zhang, Chih-Yuan Chiu, Glen Chou·· 4 小时前AI 评分32
CF-KKT:无需不安全数据,通过最优性与反事实正则化学习未知约束
Learning Unknown Constraints without Unsafe Data via Optimality and Counterfactual Regularization
AI 导读
研究者提出 Counterfactual KKT(CF-KKT)约束学习框架,结合 CIOC 的数据效率与 ICRL 的灵活性,在未知动力学下从局部最优演示中恢复未知约束,无需冒险在线探索。该方法用局部可微动力学模型施加 KKT 最优性条件,并生成奖励改进的反事实行为作为合成不可行数据。在高维机器人控制任务上,其神经网络约束表示的安全性与数据效率优于 SOTA 离线 ICRL 基线。
正文
Abstract:Learning from demonstrations (LfD) provides a framework for inferring unknown constraints from locally optimal, constraint-satisfying expert behavior. Existing approaches largely fall into two paradigms, constrained inverse optimal control (CIOC) and inverse constrained reinforcement learning (ICRL). CIOC exploits optimality conditions such as the Karush--Kuhn--Tucker (KKT) conditions but typically assumes known dynamics and structured constraint representations. Meanwhile, ICRL accommodates complex unknown constraints and unknown transition dynamics but often requires extensive online exploration, during which unsafe constraint violations may occur. In this work, we introduce Counterfactual KKT (CF-KKT), a constraint learning framework that leverages learned dynamics and locally optimal demonstrations to recover unknown constraints without requiring known dynamics or additional risky exploration, thereby combining the data efficiency and safety advantages of CIOC with the flexibility of ICRL. First, we use a locally learned differentiable dynamics model to impose KKT-inspired optimality conditions directly on the demonstrations. Second, we use the learned dynamics to generate reward-improving counterfactual behaviors near the demonstrations, revealing behaviors that would be preferable in the absence of the unknown constraint and thus providing synthetic infeasible data. When the constraint parameterization is known, the same learned-dynamics framework enables direct CIOC-based parameter recovery, and we characterize its sensitivity to dynamics misspecification. Across high-dimensional robotic control tasks, our approach learns neural constraint representations with improved safety and data efficiency relative to state-of-the-art offline ICRL baselines.
| Subjects: | Robotics (cs.RO); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.09350 [cs.RO] |
| (or arXiv:2610.09350v1 [cs.RO] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09350 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Zhouyu Zhang [view email]
[v1]
Wed, 7 Oct 2026 03:09:46 UTC (5,454 KB)
来源:arXiv:cs.LG · arxiv.org