arXiv:cs.LG· Itai Shufaro, Gal Benor, Shie Mannor·· 4 小时前AI 评分37
机制先验在序贯决策中的价值
The Value of Mechanistic Priors in Sequential Decision Making
AI 导读
研究提出"机制信息"概念,用模型推荐策略与真实最优策略间的互信息量化机制先验的价值。在渐近区间,贝叶斯遗憾随残余熵 H_mech 缩放,相比无先验基线可实现 H(μ)/H_mech 的样本复杂度降低,并给出可计算的试验前模型证书;在 burn-in 区间则建立了自信错误先验的惩罚下界。
正文
Abstract:Hybrid mechanistic models, physical priors with learned residuals, promise to reduce the data required for good decisions, but have no computable criterion to test this. We characterize the value of mechanistic priors in sequential decision-making within both asymptotic and burn-in regimes. To formalize this, we introduce the mechanistic information of a model: the mutual information between the model's recommended policy $\hat{\pi}$ and the true optimal policy $\pi^*$, bounded via a centered, occupancy-weighted bias. In the asymptotic regime (large $N$), matched bounds reveal that Bayesian regret scales with the residual entropy $H_{\mathrm{mech}}$, delivering a theoretical sample complexity reduction of $H(\mu)/H_{\mathrm{mech}}$ compared to an uninformed baseline. We further provide a computable pre-trial model certificate. Complementarily, in the clinically relevant burn-in regime (small $N$), we establish a lower bound on the penalty incurred by confidently wrong priors. We demonstrate both the asymptotic and burn-in bounds on an illustrative in-silico 5-fluorouracil (5-FU) chemotherapy plant whose structure follows published FOLFOX pharmacokinetics. The hybrid prior reduces cumulative regret by $1.79\times$ relative to standard body-surface-area dosing and $1.85\times$ relative to an uninformed learner, and remains below both under every calibration bias tested. Finally, we show that priors sensitive to distribution shift can lose half of their mechanistic information under a small distribution shift, motivating physically grounded priors for safety-critical applications.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2605.10018 [cs.LG] |
| (or arXiv:2605.10018v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2605.10018 arXiv-issued DOI via DataCite |
Submission history
From: Itai Shufaro [view email]
[v1]
Mon, 11 May 2026 05:43:53 UTC (413 KB)
[v2]
Wed, 7 Oct 2026 07:09:16 UTC (381 KB)
来源:arXiv:cs.LG · arxiv.org