跳到正文
arXiv:cs.AI· Ang Li, Yue Lin, Feifei Kou, Zhan Su, Prayag Tiwari, Wenhao Li, Shuhui Zhu, Hongyuan Zha, Baoxiang Wang·· 4 小时前AI 评分52

arXiv 论文提出可操作不确定性框架,评测 LLM 偏好推理中的行动、澄清与修复决策

Ask, Relax, or Act? Evaluating Actionable Indeterminacy in LLM Preference Reasoning

AI 导读

arXiv 论文(arXiv:2610.03102)形式化定义可操作不确定性:当所有可接受偏好下共享同一动作时执行,各选项可行但无共享动作时澄清,请求不可行时提出最小代价的约束修复。作者构建基于求解器的基准,覆盖物品分配、会议排期、公寓选择和稳定匹配,用匹配对分离决策正确性、可靠性与完全正确回答。结果显示模型常在已有合理动作时仍不必要地干预,且明确输出要求能显著提升完全正确回答并改变干预决策。

正文

View PDF HTML (experimental)

Abstract:An LLM agent can recognize uncertainty yet still choose the wrong next step: asking when action is already justified, or seeking clarification when the constraints must change. We formalize actionable indeterminacy: act when an accepted action is shared across all admissible preferences or objectives, clarify when each possibility is feasible but no action is shared, and propose a minimum-cost permitted constraint repair when the request is infeasible. We construct a solver-grounded benchmark spanning object allocation, meeting scheduling, apartment choice, and stable matching. Matched pairs retain the same source while changing whether intervention is necessary, and evaluation separates decision correctness, matched-pair reliability, and fully correct responses. Our findings reveal a recurring difficulty in recognizing when intervention is unnecessary: models can identify situations requiring clarification or repair yet still intervene when a justified action already exists. Correct decision labels also fail to guarantee usable actions, questions, or repairs. Crucially, response requirements shape not only how decisions are expressed but also which decisions are made. Making the required content explicit substantially improves fully correct responses and can change intervention decisions, even when outputs are already parseable. These findings highlight that reliable agency requires more than recognizing uncertainty: it requires intervening only when necessary and translating the chosen next step into a verifiable response.
Comments: 55 pages, 5 figures
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.03102 [cs.CL]
  (or arXiv:2610.03102v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.03102

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Ang Li [view email]
[v1] Fri, 2 Oct 2026 10:22:42 UTC (278 KB)

来源:arXiv:cs.AI · arxiv.org