跳到正文
arXiv:cs.LG· Chenyu Zhang, Rachel Luo, Boyi Li, Anjali Parashar, Marco Pavone, Apoorva Sharma·· 4 小时前AI 评分39

Careful Judge:安全高效的人类-AI 协作决策框架

Careful Judge: Safe and Efficient Human-AI Collaborative Decision Making

AI 导读

研究者提出 CARE(校准自适应修正与升级)框架,将 AI 模型与人类审核者结合,在保证人类对齐决策的同时持续从人类反馈中学习,实现更高自动化。其自适应校准模块可在每个时间步为任意修正模块提供风险控制,在驾驶、语言、机器人四个安全关键数据集上,相比基线减少 25-81% 的人类查询。CARE 具备通用、模块化特性,可适配任意黑盒 AI 模型。

正文

View PDF HTML (experimental)

Abstract:In human-AI collaborative decision making, human review can prevent unsafe AI decisions, but each human judgment is costly. Treating human intervention after AI abstention as a one-off fallback misses the opportunity to improve future AI decisions for greater automation, yet AI adaptively learning from selectively queried human feedback breaks safety guardrails calibrated for old models. We approach this challenge with CARE---calibrated adaptive rectification and escalation---an end-to-end pipeline that combines AI models and human reviewers to guarantee safe, human-aligned decisions, while continuously learning from human feedback to achieve greater automation with fewer human queries. CARE is principled, general, modular, and works with any black-box AI model. Our novel adaptive calibration module guarantees risk control at every time step for any rectification module. We further show how CARE improves query efficiency when the AI model is well trained and the human-AI misalignment has a clear structure. Experiments on four safety-critical real-world datasets spanning driving, language, and robotics demonstrate that CARE achieves human-aligned decisions while reducing human queries by 25-81% relative to baselines.
Subjects: Machine Learning (stat.ML); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2610.09043 [stat.ML]
  (or arXiv:2610.09043v1 [stat.ML] for this version)
  https://doi.org/10.48550/arXiv.2610.09043

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Chenyu Zhang [view email]
[v1] Tue, 6 Oct 2026 19:44:15 UTC (7,246 KB)

来源:arXiv:cs.LG · arxiv.org