arXiv:cs.AI(全量分类)· Yifan Zhang, Liang Zheng·· 5 小时前AI 评分32
MC-WM:面向受控世界模型适配的校准风险路由
Calibration-risk routing for controlled world-model adaptation
AI 导读
研究者提出 Model-Corrected World Model(MC-WM),将初始目标数据划分为互不重叠的拟合、选择与校准三部分,并按标准化校准风险更低者选择模型族。该方法用学习到的置信信号与确定性有效性谓词加权一步想象策略更新,且不改写物理奖励。在三个受控 MuJoCo 动力学偏移上共完成 541 次执行,覆盖 540 个唯一报告运行单元。
正文
Abstract:Model-based reinforcement learning (MBRL) can exploit simulated experience, but a simulator-to-target shift creates a model-selection problem: correcting the simulator and fitting the target directly can each fail under limited target data. We introduce the Model-Corrected World Model (MC-WM), which separates initial target data into disjoint fit, selection, and calibration partitions and deploys the family with lower standardized calibration risk. A learned confidence signal and deterministic validity predicates weight one-step imagined policy updates without rewriting physical rewards. We evaluate 540 unique reported run cells across three controlled Multi-Joint dynamics with Contact (MuJoCo) shifts; one exact-routing cell was repeated after a pre-deployment artifact gate, giving 541 completed executions.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.01001 [cs.AI] |
| (or arXiv:2610.01001v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.01001 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yifan Zhang [view email]
[v1]
Thu, 1 Oct 2026 03:44:47 UTC (91 KB)
来源:arXiv:cs.AI(全量分类) · arxiv.org