跳到正文
arXiv:cs.AI· Chubin Zhang, Zhenglin Wan, Xingrui Yu, Jingxuan Wu, Qi Wen, Pengfei Zhou, Wangbo Zhao, Ivor Tsang·· 6 小时前AI 评分61

Calibration Is Not Control:用干预价值监督 LLM 智能体

Calibration Is Not Control: Intervention Value for LLM-Agent Oversight

AI 导读

arXiv 论文(NeurIPS 2026,arXiv:2606.21399)提出运行时监督不应只依赖校准的失败分数阈值,因为失败风险相同的状态在干预收益上可能不同,而严格递增的再校准无法区分二者。

正文

View PDF HTML (experimental)

Abstract:Runtime oversight often intervenes when an LLM agent's calibrated failure score crosses a threshold. Yet states with the same failure risk can differ in whether intervention helps. Strictly increasing recalibration preserves the threshold policy class and cannot recover this distinction. We formalize when a summary is sufficient for intervention decisions and the utility lost when it is not. We evaluate the consequences by replaying agent prefixes and executing alternative actions from the same state. On ALFWorld, holding features, estimator, and router fixed while changing the supervision target from failure to intervention utility lowers regret from 0.51 to 0.09; the gain replicates on a second suite of mid-episode prefixes. A deployable intervention-trained scalar also beats the failure-score threshold rule selected on test outcomes. Online, on 300 unseen tasks with a fixed stronger-model handoff, a frozen prefix-feature controller improves utility over failure-triggered routing, handing off less often (35% vs 48%) and succeeding more often (45% vs 37%). Gains depend on intervention value and are small on two reasoning benchmarks. Oversight signals should be evaluated by the decisions they support alongside their predictive quality. Code is available at this https URL.
Comments: NeurIPS 2026
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2606.21399 [cs.AI]
  (or arXiv:2606.21399v2 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2606.21399

arXiv-issued DOI via DataCite

Submission history

From: Chubin Zhang [view email]
[v1] Fri, 19 Jun 2026 13:08:17 UTC (3,542 KB)
[v2] Tue, 6 Oct 2026 12:18:59 UTC (1,583 KB)

来源:arXiv:cs.AI · arxiv.org