arXiv:cs.LG· Heewon Park, Somin Im, Minhae Kwon·· 4 小时前AI 评分54
SAG:面向长程决策 LLM Agent 的成本感知选择性 Critique 框架
Selective Critique for Cost-Aware LLM Agents in Long-Horizon Decision Making
AI 导读
论文提出 SAG(Self-improving Agent with Gated critique),将 critique 调用建模为逐步决策问题,用动作级歧义信号(全局熵与 top-2 margin)构成免训练门控,近似 critique 的信息价值(VoI),仅在预期收益超过成本时调用反馈,并通过在线自举让 actor 逐步内化 critic 辅助行为。
正文
Abstract:Improving the reliability of large language model (LLM) agents in long-horizon decision-making remains a key challenge. When deployed as autonomous agents interacting with complex environments, early mistakes can propagate through trajectories and cause cascading failures. Recent approaches improve reliability by incorporating external critique or deliberation, but invoking these mechanisms at every step substantially increases token consumption and latency, limiting practical deployment. We propose SAG (Self-improving Agent with Gated critique), a cost-aware framework that formulates critique invocation as a step-wise decision problem during long-horizon interaction. SAG introduces a lightweight, training-free gating mechanism that estimates the utility of critique using action-level ambiguity signals--global entropy and local top-2 margin--computed over admissible actions. From a decision-theoretic perspective, this mechanism approximates the Value of Information (VoI) of critique, enabling the agent to selectively allocate expensive feedback only when its expected benefit justifies the cost. SAG further incorporates online bootstrapped self-improvement, allowing the actor to internalize critic-assisted behaviors and progressively reduce reliance on critique. Across three long-horizon interactive benchmarks and multiple backbone models, SAG substantially improves the performance-cost trade-off compared with both no-critique and always-on critique agents. On ALFWorld, SAG increases task success from 24.6% to 78.4% while maintaining a token budget comparable to ReAct, yielding a $3.1\times$ improvement in normalized token efficiency. Moreover, a 7B actor with a lightweight 3B critic achieves performance comparable to a 14B actor without critique, showing that selective critique can recover most of the reliability benefits of deliberation while dramatically reducing inference cost.
| Comments: | Accepted to NeurIPS 2026 |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.07335 [cs.LG] |
| (or arXiv:2610.07335v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07335 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Heewon Park [view email]
[v1]
Mon, 5 Oct 2026 20:09:48 UTC (3,040 KB)
来源:arXiv:cs.LG · arxiv.org