跳到正文
原文
arXiv:cs.AI(全量分类)· Yezhou Cheng, Runjia Du, Zeming Liu, Hang Lyu, Zehua Yang, Bojun Lin·· 5 小时前AI 评分36

GAVA:文本具身智能体中用户纠正的接地仲裁

Knowing When to Yield: Grounded Arbitration of User Corrections in Text-Based Embodied Agents

AI 导读

GAVA 提出"接地纠正仲裁"框架,让具身智能体在用户纠正可能有误时,在接受、拒绝、检查世界和询问说话者之间做选择,通过观察受限证据、合法探针和一步期望损失规则实现。

正文

View PDF HTML (experimental)

Abstract:How should an embodied agent respond when a person's correction may be wrong? We formulate grounded correction arbitration as a choice among accepting, rejecting, inspecting the world, and asking the speaker. GAVA implements this interface with observation-bounded evidence, legal probes, and a one-step expected-loss rule. In text-only ALFWorld, 162 checkpoints produce 972 paired true and false interventions. Complete local inspections give GAVA and always verify 100 percent correction accuracy, establishing the evidence contract rather than a comparative advantage. In same-episode execution, GAVA reduces interaction cost against always verify but ties a cost threshold under a perfect speaker. An exploratory training-only object-location prior lowers interaction and declared joint cost on 340 unseen scenarios by 0.490 and 0.420 relative to uniform GAVA. After freezing the policy, costs, baselines, and multiplicity plan, the gains replicate on 77 non-overlapping seen checkpoints, covering 308 scenarios: 0.595 and 0.517, with both 95 percent checkpoint-bootstrap confidence intervals excluding zero. Joint cost also improves over an identical-prior fixed policy, while the matched calibrated no-VOI comparison remains inconclusive. Semantic GAVA makes four factual errors in each cohort, corresponding to 98.8 percent and 98.7 percent accuracy, and all methods complete every task. Results support selective information gathering with semantic priors under declared costs, but do not establish a general advantage of environmental value of information over clarification. The study uses normalized claims, complete symbolic observations, and controlled speakers; it evaluates neither human participants, visual input, nor physical robots.
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.00282 [cs.AI]
  (or arXiv:2610.00282v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.00282

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Yezhou Cheng [view email]
[v1] Fri, 25 Sep 2026 02:28:24 UTC (74 KB)

来源:arXiv:cs.AI(全量分类) · arxiv.org