跳到正文
arXiv:cs.LG· Yannis Montreuil, Axel Carlier, Lai Xing Ng, Wei Tsang Ooi·· 4 小时前AI 评分28

ORUCB:在线延迟决策中学习修正专家答案

A Query Is Not a Commitment: Learning to Correct Expert Answers in Online Deferral

AI 导读

针对在线延迟决策中专家答案需先修正再使用的问题,研究者提出 ORUCB,将共享响应与专家专属多项式响应联合建模,并用累积响应学习误差上界校准置信度加权风险回归与探索。在固定问题参数下,该算法实现 O(√T log(T+1)) 的高概率伪遗憾。在四个测试流上,所选三次策略的费用含成本低于七个直接采用原始答案的基线。

正文

View PDF HTML (experimental)

Abstract:An inaccurate expert can still provide useful information after correction. We study online learning to defer in which the learner chooses an expert and fixes a correction function before purchasing its answer, then applies that function to the answer received. The difficulty is that observed losses reflect both expert quality and an unfinished correction: early errors can discourage queries that would be valuable after learning. We propose ORUCB, which pools shared and expert-specific polynomial responses. A bound on cumulative response-learning error calibrates confidence-weighted risk regression and exploration, allowing the router to account for this error when deciding which answers to buy. Under bounded residuals and disagreements, a fixed feasible model of optimal responses, and linear models of free and optimal queried risk, the calibrated algorithm achieves high-probability pseudo-regret $O(\sqrt T\log(T+1))$ over $T$ rounds for fixed problem parameters. The guarantee permits singular answer distributions and misspecified shared responses; optimality is relative to the bounded response class. On four test streams, the selected cubic policy has lower fee-inclusive cost than seven baselines that deploy answers unchanged. Comparisons with a common correction learner examine routing, while six-price comparisons measure cost and query rates.
Subjects: Machine Learning (stat.ML); Machine Learning (cs.LG)
Cite as: arXiv:2610.07084 [stat.ML]
  (or arXiv:2610.07084v1 [stat.ML] for this version)
  https://doi.org/10.48550/arXiv.2610.07084

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Yannis Montreuil [view email]
[v1] Mon, 5 Oct 2026 11:55:34 UTC (952 KB)

来源:arXiv:cs.LG · arxiv.org